CaseAdvancedDesigning for Uncertainty & Trust / UX for uncertainty and confidence display / #17

What research would you run to test whether your uncertainty display works?

LEAD the behavioral signal that would have told Naomi's team the truth, weeks before the disputes did

Trellis Home matches homeowners with contractors for repair jobs and shows a Match Confidence tag, High, Medium, or Low, on every recommendation. Naomi Castellano is the UX researcher who joined the team to find out whether that tag actually changes what people do, not just whether they notice it.

The direct answer
Don't ask people if they noticed the tag. Watch whether they act differently depending on it: log how much a homeowner shops around and how the job actually turns out, split by confidence tier, not averaged together. If a Low tag doesn't produce more comparing and a High tag doesn't produce faster, calmer booking, the tag isn't working, no matter what a survey says.
Do this, in order
  1. Log real behavior by confidence tier, not just whether the tag was shown.Why: an impression count tells you the tag appeared; it tells you nothing about whether anyone changed course because of it.
  2. Watch a small group of real homeowners live, don't just ask them afterward.Why: people say they noticed things they never actually acted on; watching catches the gap a survey can't.
  3. Test the tag's wording against at least one alternative, measured on behavior, not preference.Why: a wording people "like" in a survey can still fail to change a single booking decision.
  4. Set a real action for each finding before you start, not after.Why: research nobody planned to act on becomes a slide deck, not a decision.
  5. Don't run this whole study on every minor tag tweak going forward.Why: the full behavioral study earns its cost for a real design question, not for a font-size change.

How to answer this, stage by stage

The interviewer isn't checking whether you can name a research method. They're checking whether you'd trust a survey that says "yes, I saw it," over behavior that says otherwise.

Move 1
Frame the one question that matters
Say it like this
"I'll answer this for Trellis Home's Match Confidence tag. The real question isn't 'do people see it,' it's 'does it change what they actually do next.'"
Why this works
Reframes "test whether it works" away from a satisfaction score before any method gets named.
Move 2
Name the method you're running
Say it like this
"I'll use LEAD. Link to the real outcome, find the early signal, watch how the metric could be gamed, and decide what action each result triggers."
Why this works
Shows the interviewer this is a designed research plan, not an ad hoc list of things to check.
Move 3
Say what "success" actually means here
Say it like this
"Success isn't 'people notice the tag.' It's homeowners shopping around more on Low-confidence matches and booking with less hesitation on High-confidence ones."
Why this works
Names the real behavioral outcome, so the research design has something concrete to actually measure.
Move 4
Give the research design in one breath
Say it like this
"Log booking behavior by tier, watch a small group live in moderated sessions, and A/B test the wording against a real behavioral metric, not a preference question."
Why this works
This is the direct answer, stated as three concrete methods, before any story tries to sell it.
Move 5
Show why a survey alone would have lied
Say it like this
"Trellis Home's original satisfaction survey said 'yes, I noticed the tag' from 78% of respondents. Actual booking behavior barely differed between confidence tiers at all."
Why this works
Turns "self-report can mislead you" into one specific, checkable gap.
Move 6
State the action at each finding
Say it like this
"If behavior genuinely diverges by tier, ship it. If people are engaged but act the same either way, redesign the wording. If it backfires, pull it back to a small test group."
Why this works
Shows the research has a real decision waiting at the end, not just a report.
Move 7
Close with the real test
Say it like this
"Don't test whether the tag was seen. Test whether it moved someone's hand toward comparing one more contractor, or toward booking with less hesitation."
Why this works
Restates the direct answer, ready for a live follow-up.

Let's learn

Say a marketplace builds a feature that tells homeowners how much to trust its top contractor pick. How would you actually know it works?

Trellis Home's Match Confidence tag appears on every recommended contractor: High, Medium, or Low, based on how much history the algorithm has on that contractor and how well they match the specific job. Before the tag existed, homeowners booked the top-ranked contractor directly about 71% of the time, with no signal at all about how sure the match really was.

Hand sketched flow diagram titled The research process step by step. Five boxes: recruit real homeowners, show live confidence tags, log what they actually do highlighted, compare tiers not averages, decide ship redesign pull back.
The middle box is the one most teams skip, watching what people do, not just asking them about it afterward.

After the tag shipped, the team's only proof it worked was a satisfaction survey: 78% of respondents said they'd noticed the tag, and Trellis Home called that a launch success.

Here's the turn: noticing a tag and acting on it are not the same fact, and nobody had ever checked whether the second one was actually happening. The real risk wasn't that the tag failed. It's that nobody could tell the difference between it working and it just sitting there.

Post-job dispute rate, Low-confidence vs. High-confidence matches, before and after a behavior-tested redesign
25% 12.5 0 22% 14% Before redesign 11% 13% After redesign
Red is Low-confidence matches, blue is High-confidence. Before the redesign, a Low tag barely changed the dispute rate at all. After, it nearly closed the gap.

At its worst, an untested confidence display can quietly do nothing while everyone believes it's doing its job, right up until leadership asks for proof and the team realizes they never built a way to check.

The decision I would take back Nobody logged what a homeowner did after seeing a given confidence tier, whether they booked immediately, browsed more contractors, or abandoned the flow. That made sense at launch, when the tag was brand new and logging every downstream action felt premature. It stopped making sense once leadership wanted proof the feature was worth keeping, and the team realized there was no history to check it against.

What I would leave alone: the underlying match algorithm itself doesn't need this scrutiny, its accuracy was already tracked separately. This research is about whether the display of that accuracy changes behavior, a different question entirely.

The lesson: a feature that shows people something is only as good as your ability to prove it changed what they did next. Without that, "it works" is a guess wearing a launch announcement.

Now here is the same thing as a story

The short version above is what you'd say out loud in an interview. Read this one for how the real gap actually surfaced.

The research lab at Trellis Home empties out by 6, except on Tuesdays, when Naomi Castellano runs live sessions with homeowners testing new features. Six weeks into her first quarter there, in a routine planning meeting, she asked a simple question: how do we know the Match Confidence tag is actually changing anyone's behavior, versus just sitting there?

Nobody on the existing team had a real answer. They had impression counts, showing the tag rendered correctly on-screen. They had the satisfaction survey, showing 78% of respondents recalled seeing it. Neither one said anything about what a homeowner actually did next.

Knowledge spark: what's the difference between a leading and a lagging metric here? A lagging metric, like a dispute rate, only shows up weeks after a job is booked and finished. A leading metric, like how many extra contractor profiles someone views before booking, shows up the same day, and predicts the lagging one weeks in advance, if the design is actually working.

Naomi's team recruited 40 real homeowners with an active repair need and watched them use a live prototype in moderated sessions, rather than asking afterward whether they'd "noticed" anything. Watching mattered: several participants said, unprompted, that they'd seen the tag, then booked the top match anyway without opening its reasoning, exactly the gap a survey alone would have missed entirely.

Hand sketched metaphor scene titled How the research gets gamed. Left, a document icon labeled Survey says, caption yes I noticed the tag. Right, a person icon labeled Real booking, caption books the same way regardless.
The survey and the actual booking told two different stories. Only one of them was true.

Alongside the moderated sessions, the team logged real behavior across the full platform, split by confidence tier instead of averaged together. In week 1 of the study, homeowners viewed extra contractor profiles at nearly the same rate regardless of tag, around 30% for both Low and High confidence matches. That flat line was the actual finding: the tag wasn't changing anything yet.

Hand sketched comparison titled Two clocks, ringing at different times. Left panel, a gauge icon labeled Leading signal, caption rings in week 3. Right panel, a gauge icon labeled Lagging signal, caption rings in week 9.
The leading signal would have rung in week 3. The dispute-rate data wouldn't have caught up until week 9.
The team almost shipped a feature that felt successful and did nothing, because the only number they'd ever checked was whether people said they'd seen it.

The team ran an A/B test on the tag's wording, comparing the original generic label against a version naming the specific reason, "Only 2 similar jobs completed nearby" instead of just "Low confidence." Measured against real behavior, not preference, the specific-reason version produced a real, growing gap between tiers within about 8 weeks.

Hand sketched decision tree titled What did the behavioral data actually show. Root: what did the behavioral data show. Branches: real divergence by tier leads to ship as is, engaged but no behavior change leads to redesign the wording, diverges the wrong way leads to pull back investigate.
Three possible findings, three different actions. Nobody had written this down before the study started.

LEAD, mapped onto this one research planNot a lecture on research methods in general. LEAD is what tells you which number would have caught the gap first.

L
Link. The real outcome at stake.
Not "did people notice the tag," but whether homeowners' actual booking behavior became appropriately matched to how sure the algorithm really was.
Grounds the whole study in the thing that actually costs something: mismatched jobs and disputes.
E
Early signal. The number that moves first.
The rate of homeowners viewing extra contractor profiles after a Low-confidence match, split by tier. It diverged from the High-confidence rate weeks before the dispute rate ever changed.
This is the hardest step and the answer to the question: real behavior, split by tier, is the leading signal a satisfaction survey can never be.
A
Abuse. How this metric gets gamed.
A raw click-through count on the "why this match" panel can rise from curiosity alone, without ever changing whether someone books differently.
Explains why behavior, not engagement, has to be the thing actually measured.
D
Decision. What you'd do at each threshold.
Real divergence by tier: ship it. Engagement with no behavior change: redesign the wording. Divergence the wrong way: pull back and investigate.
Turns the research into a decision waiting to happen, not a report that sits on a shelf.
Hand sketched labeled parts diagram titled What's inside the research plan. Center icon document labeled Research Plan, with four callouts: behavior log by tier, moderated sessions, a wording A B test, the real outcome metric.
Four parts, and none of them is a satisfaction survey question on its own.
Share of homeowners viewing 2 or more extra contractor profiles, by match confidence tier, week by week
60% 30 0 Week 1 Week 8 Low confidence High confidence
The lines split apart weeks before the redesigned dispute-rate numbers confirmed it. This gap is what a satisfaction survey could never have shown.

The recap, one line per letter: link is genuinely calibrated booking behavior, not tag recall, early signal is the browse-more rate split by tier, abuse is mistaking curiosity clicks for real behavior change, and decision is a pre-set action for each of the three possible findings.

And if you want to be sure it really works, try it somewhere elseSame four letters, a food-delivery app's ETA confidence display instead of a contractor marketplace. A different lagging outcome proves the same early signal works.

Corvid Eats shows a delivery-time estimate with a small "less certain, traffic is heavy" note attached when its ETA model has lower confidence, and product researcher Bertrand Achebe wants to know whether that note actually changes anything, or just makes an occasionally-late delivery feel more forgivable after the fact.

Mapped onto LEAD: link is whether customers' actual patience and reorder behavior improve when expectations are honestly set, not just whether complaint tickets read as less angry. Early signal is the rate of customers who check the live map tracker again within 5 minutes of an uncertain-ETA order, versus a confident one, since checking more suggests the note registered as real information, not decoration. Abuse is a naive version of this research counting total app opens as "engagement," when a customer anxiously refreshing the map is actually a sign of unresolved uncertainty, the opposite of the note working well. Decision: if checking behavior on uncertain orders drops toward the confident-order baseline over a few weeks, the note is genuinely reassuring people appropriately; if it stays elevated indefinitely, the wording needs a concrete next step, like "we'll text you if this changes," not just an acknowledgment of doubt.

Hand sketched icon list titled Signs your research is measuring the wrong thing. Four items: only asks if they noticed it, no control group at all, never checks the real outcome, stops measuring the day it ships.
All four of these were true of Trellis Home's original launch survey. None of them were true of the redesigned study.

The old decision at Corvid Eats wasn't quite Trellis Home's absent-logging gap, it was closer to a packaging choice: the ETA note was written once at launch and never revisited, treated as a finished, static piece of copy rather than something whose real effect needed checking against actual customer behavior over time, the same way Trellis Home once assumed its own tag was working simply because it had shipped.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "log real behavior by confidence tier, not just impressions, and set an action for each possible finding before you start," and stop.
Cost: there's no budget for a full moderated study right now. Say so, and start with the free part: split existing behavioral logs by confidence tier retroactively; that alone often reveals whether a gap even exists before spending on new research.
The model gets better, for real: if the underlying match algorithm improves so much that Low-confidence tags become rare, the same behavioral-logging discipline still matters, it just has less to show on a shrinking slice of traffic.

Where people run it wrong.
They treat a satisfaction survey's "yes, I noticed it" as proof the feature works.
They count engagement, like clicks or opens, without checking whether it changed a real decision.
They launch research with no pre-agreed action, so every finding becomes a debate instead of a decision.

How to use it live. When someone asks how you'd test an uncertainty display, ask back: what would someone do differently, in their own hands, if this were actually working? Design the study around answering that, not around whether people remember seeing a label.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a "how would you test this metric or display" question, and what does each letter do?
Tap to flip
ANSWER
LEAD: link, early signal, abuse, decision. It finds the real behavioral signal that moves before the lagging outcome does, and pre-commits an action to each finding.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Naomi Castellano, the UX researcher at Trellis Home who asked whether the Match Confidence tag actually changed behavior, not just whether people noticed it.
3 · THE LINK
What's the real outcome this research is actually trying to protect?
Tap to flip
ANSWER
Whether homeowners' booking behavior becomes appropriately calibrated to how sure the match really is, not whether they recall seeing a tag.
4 · THE EARLY SIGNAL
What's the leading indicator this answer tracks, and why not just wait for the dispute rate?
Tap to flip
ANSWER
The share of homeowners viewing 2 or more extra contractor profiles, split by confidence tier. It diverged weeks before the dispute rate ever moved.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Never logging what a homeowner actually did after seeing a given confidence tier, leaving no history to check whether the tag was doing its job.
6 · THE NUMBER
Fill in the blank: in week 1, both confidence tiers showed about ___ percent of homeowners viewing extra contractor profiles, meaning the tag wasn't changing anything yet.
Tap to flip
ANSWER
About 30 percent. A flat, matching rate across tiers is exactly what "the tag isn't working yet" looks like in real data.
7 · THE REPLAY
Same launch, same original tag, but the team runs LEAD's behavioral study from day one instead of trusting the satisfaction survey. What changes?
Tap to flip
ANSWER
They catch the flat browse-rate gap in week 1 or 2, redesign the wording to name a specific reason, and see real divergence by week 8, instead of finding out the hard way from dispute data months later.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the early signal there?
Tap to flip
ANSWER
Corvid Eats' food-delivery ETA confidence note. The early signal is how often customers re-check the live map tracker within 5 minutes of an uncertain-ETA order.

Check yourself Score: 0 / 0

Multiple choice
1. Why did Naomi's team distrust the original 78% "noticed the tag" survey result?
  • A. The survey sample size was too small to matter.
  • B. Noticing a tag and actually changing behavior because of it are two different facts, and only one had ever been measured.
  • C. Respondents were suspected of lying about seeing the tag.
  • D. The survey was run by a different department than the design team.
Show hint
Look at "here's the turn."
Show answer
B. A recalled tag and a changed decision aren't the same event, and moderated sessions showed people doing exactly this: noticing it, then acting the same anyway.
True or false
2. True or false: counting how many people clicked into the "why this match" reasoning panel is, by itself, enough proof the confidence display is working.
  • True
  • False
Show hint
Look at the abuse step.
Show answer
False. Clicks can happen out of curiosity without changing a single booking decision; only a real behavior difference by tier counts as proof.
Fill in the blank
3. Fill in the blank: after the wording redesign, the Low-confidence dispute rate fell from 22 percent to about ___ percent, close to the High-confidence rate of 13 percent.
Show hint
Look at the grouped bar chart in Section 1.
Show answer
11 percent. The two tiers nearly converged, which is what "the tag is finally doing its job" looks like in the lagging outcome.
Short answer, apply it yourself
4. Think of an app that shows you some kind of confidence or trust signal, like a seller rating or a "verified" badge. How would you actually test whether it changes what you do, rather than just whether you've seen it?
Show hint
Think about what real action would look different if the signal were actually working.
Show answer
Model answer: Most people have never thought about this distinction before, which is exactly why so many "trust signals" in real products go untested.
Short answer, why no middle setting
5. Why wasn't "just ask people in a follow-up survey whether the tag helped them" a sufficient research plan on its own?
Show hint
Look at how the metric could be gamed.
Show answer
Model answer: Self-report captures what people believe about their own behavior, which moderated sessions showed doesn't always match what they actually did.
Short answer, where it wouldn't matter
6. Name a part of Trellis Home's system where this full behavioral research plan genuinely isn't needed.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The underlying match algorithm's own accuracy tracking. That's a separate, already-measured question from whether the display of that accuracy changes behavior.
Before you close the answer
Why this works
Tests whether you'll accept a self-reported "yes, I saw it" as proof a design decision worked, or insist on watching real behavior split by the exact condition you're trying to change.
Follow-up traps
"Isn't a moderated session with 40 people too small to trust?" Response: it's not meant to stand alone; it explains the "why" behind the platform-wide behavioral logs, which cover every real booking, not just 40 sessions.

"What if behavior differs by tier, but the actual jobs don't turn out any better?" Response: then the dispute-rate data, the true lagging outcome, would have to be checked too; a behavior change that doesn't improve real job outcomes isn't the finish line either.
If pressed
The wording A/B test ran for a full 8 weeks specifically because early results in week 2 looked like a real gap, but it was mostly driven by one unusually busy zip code; waiting out the full period is what kept the team from shipping a fix to noise instead of a fix to the actual problem.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more