CaseAdvancedDesigning for Uncertainty & Trust / Onboarding users to probabilistic products / #19

Describe how you would measure whether onboarding taught the right mental model.

LEAD the product is Farrowgate, a secondhand-goods marketplace, and ValueMark is the AI that suggests a listing price

Farrowgate is a marketplace for buying and selling used goods between neighbors. ValueMark suggests a listing price from photos and a condition description. Bettina Voss sells off her late grandmother's household items on Farrowgate most weekends.

The direct answer
Don't wait for the dispute rate to tell you onboarding failed. Right after onboarding, and again through a seller's first ten listings, show them one sample item and ask them to guess whether ValueMark's suggested price will land above or below their own gut number, before revealing it. Track that prediction-accuracy rate over time. A seller with the right mental model gets better at guessing. A seller who never formed one stays stuck near a coin flip, or stops guessing at all.
Do this, in order
  1. Track prediction accuracy on a direction-and-size guess, not acceptance rate.Why: accepting a number tells you nothing about whether the seller understands why it's there.
  2. Measure it as a trend across a seller's first ten listings, not a single score.Why: a healthy mental model shows a learning curve; a broken one plateaus early or never engages.
  3. Weight the check questions toward cases where the model diverges from naive intuition.Why: a seller who always guesses "same as my number" can look accurate without ever really predicting anything.
  4. Pair the leading indicator with the lagging outcome it's supposed to predict.Why: a metric that doesn't connect to real disputes or seller regret later is just a number nobody should act on.
  5. Trigger a short re-explainer only for sellers stuck near chance by listing five.Why: nagging sellers who already understand it trains everyone to ignore the check.

How to answer this, stage by stage

Nobody's grading whether you can name a metric. They're grading whether the metric you pick would have caught the problem before the damage was done.

Stage 1
Scope it to one real onboarding moment
Say it like this
"I'll answer this for Farrowgate's ValueMark feature, and for what a new seller like Bettina is actually taught to believe about that suggested price."
Why this works
Anchors "mental model" in one specific belief, not a vague notion of understanding.
Stage 2
Say your structure out loud
Say it like this
"I'll use LEAD. Link, the real business outcome. Early signal, what moves first. Abuse, how it gets gamed. Decision, what I'd actually do at each threshold."
Why this works
Signals a method built specifically for metric questions, not a borrowed framework.
Stage 3
Name the outcome that actually matters
Say it like this
"The real outcome is fewer disputes and fewer sellers who feel cheated after underselling something rare, because both of those drain trust in the marketplace itself."
Why this works
Ties the metric to something the business genuinely cares about, not a vanity engagement number.
Stage 4
Give the leading signal
Say it like this
"Right after onboarding, show a sample item and ask the seller to guess: will ValueMark's price be higher, lower, or about the same as their gut number? Track how often they're right, across their first ten listings."
Why this works
Measures the actual belief, not a behavior that could mean several different beliefs.
Stage 5
Name how it gets gamed
Say it like this
"If most check items are ones ValueMark handles the way anyone would guess, a seller can look accurate without ever really predicting anything. I'd weight the checks toward the cases where the model actually surprises people."
Why this works
Shows you've thought about how the metric fails, not just what it measures on a good day.
Stage 6
Say what you'd actually do at each threshold
Say it like this
"If a seller is still guessing near a coin flip by listing five, I'd show them one worked example of a real miss. If they're already accurate, I'd stop showing the check entirely."
Why this works
A metric nobody acts on is a dashboard decoration. This makes the number actionable.
Stage 7
Name the rejected alternative
Say it like this
"I considered just tracking how often sellers accept the suggested price unchanged. I rejected it, since a high acceptance rate looks identical whether someone understands the tool or is just clicking through."
Why this works
Proves this was a real choice, and names exactly why the more obvious metric fails.
Stage 8
Close on the one line
Say it like this
"Measure whether they can predict it, not whether they click it. A right mental model shows up as a learning curve, weeks before a single dispute ever gets filed."
Why this works
Restates the direct answer once more, so it's the last thing an interviewer hears.

Let's learn

Here is a question worth sitting with: what would it look like, weeks before anyone complains, if onboarding had quietly taught someone the wrong idea?

Farrowgate is a marketplace for secondhand household goods. ValueMark suggests a listing price for anything a seller photographs and describes.

Knowledge spark: what's a leading indicator? A number that moves before the outcome you actually care about does. It's useful precisely because it gives you weeks of warning instead of a dispute that's already happened.

Before onboarding taught anything about ValueMark, sellers guessed prices from nothing but instinct, often wildly off in both directions. After onboarding, most sellers accept ValueMark's suggestion outright, about 74% of the time, unedited.

Prediction accuracy on the direction-and-size check, by listing number
100% 50% 0 calibrated stuck at chance Listing 1 Listing 5 Listing 10
One line climbs. One line never moves off a coin flip. Both groups accept ValueMark's price at almost the same rate, so acceptance alone would never have told them apart.

At its worst: a seller who never formed a real mental model finally hits a rare, oddly-shaped item, a hand-carved chest with sentimental but no obvious resale value, accepts a low suggested price without question, and only feels the loss weeks later scrolling similar items that sold for triple.

The decision that mattered Measure prediction, not acceptance. A seller who can guess ValueMark's direction correctly, more often over time, has the mental model onboarding was supposed to build. A seller who accepts every number without ever engaging in a real guess might have the same acceptance rate, and a completely different, riskier understanding underneath it.

What I would leave alone: a seller whose prediction accuracy is already high by listing three doesn't need the check repeated on every future listing. Once the mental model is there, more checking is just friction with nothing left to measure.

A high acceptance rate and a right mental model can look exactly the same on a dashboard, right up until the one item where they aren't the same thing at all.

The lesson: the metric that would have caught this isn't a bigger number further downstream. It's a smaller, earlier one that asks the seller to commit to a guess before they see the answer.

Now here is the same thing as a story

The short version above is what you'd say defending this metric in a product review. Read this one for how the gap actually showed up for Bettina.

Bettina lists most of what's left of her late grandmother's house on Sunday afternoons, one box at a time, coffee going cold on the counter.

Hand sketched labeled parts diagram titled The comprehension check, close up. Center document icon labeled One sample item, with four callouts: guess up or down, guess by how much, tap to reveal, one screen ten seconds.
The whole check fits in ten seconds. It asks for a guess, not a click, and that's the entire difference.

In her first week, onboarding walked her through exactly what ValueMark is: a suggestion built from photos and condition, not an appraisal. She nodded along, the way anyone does through an onboarding screen, and moved on to her first listing.

Hand sketched flow diagram titled Bettina's first ten listings. Four steps: lists item one, guesses wrong, guesses right more, trusts calibrated.
Bettina's line on the chart above is this flow, in numbers. She got it wrong first, and that's exactly what made the later guesses meaningful.

By her fifth listing, Bettina's prediction accuracy had climbed to 68%. She'd learned that ValueMark tends to price sentimental-but-common items lower than instinct suggests, and rare, well-preserved items higher. That's the mental model onboarding was actually trying to build.

Hand sketched comparison diagram titled Two mental models. Left panel, a box icon labeled Accepts blindly, caption never predicts just clicks. Right panel, a gauge icon labeled Predicts then checks, caption guesses direction first.
Both sellers click accept about as often. Only one of them is actually thinking about the number before they do.

Another seller in the same cohort, someone who never engaged with the guess-first check at all, kept an identical 74% acceptance rate the entire time. On paper, that looked just as healthy as Bettina's.

Hand sketched quadrant titled Sorting sellers by mental model health. Axes prediction accuracy and dispute rate. Calibrated sellers sit at climbing accuracy and low dispute rate. Blind accepters sit at near chance accuracy and high dispute rate. New sellers in week one sit in the middle.
Bettina and the other seller start in nearly the same spot. Only the prediction check shows which direction each of them is actually moving.
Hand sketched timeline titled When each metric would have rung. Three milestones: listing 3 prediction rate stalls, listing 7 still stalled unnoticed, week 6 undersell regret complaint.
The gap between the first mark and the last one is four weeks of warning a dispute-rate metric alone would never have given anyone.

Six weeks later, that other seller listed a rare item and accepted a suggested price well under what a specialist later told them it was worth. The complaint that reached Farrowgate's support team blamed the tool. The real story was a mental model that had never actually formed, sitting quietly behind a completely normal-looking acceptance number the whole time.

Hand sketched metaphor scene titled How the check gets gamed. Left, a document icon labeled CHECK PASSES, caption always guesses same. Right, a person icon labeled STILL BLIND, caption no real prediction.
A seller who always guesses "same as my number" can pass the check without ever really predicting anything, which is exactly why the check items have to include cases where the model actually surprises people.

The old measurement asked only whether sellers used ValueMark. The new one asks whether they understand it, weeks before an unusual item ever puts that understanding to the test.

We built the acceptance-rate dashboard first because it was the easiest number to pull and it always looked good. It took one specialist's appraisal, arriving after the fact, to see that a healthy-looking acceptance rate and a real mental model were never actually the same measurement.

LEAD, in one screenNot the dispute rate three months later. The guess a seller makes before they ever see the answer.

L
Link. The real outcome.
Fewer disputes and fewer sellers who feel cheated after underselling something rare.
Ties the whole metric to something the business genuinely loses money and trust over.
E
Early signal. The direction-and-size guess.
Prediction accuracy on a sample item, checked before onboarding shows the real price, tracked across a seller's first ten listings.
The hardest step, and the actual answer to the question being asked.
A
Abuse. How it gets gamed.
Guessing "same as my own number" every time can look accurate on items where ValueMark agrees with naive intuition anyway.
Names the way the metric could be satisfied without any real comprehension underneath it.
D
Decision. What happens at each threshold.
Stuck near chance by listing five: show one real miss example. Already accurate: stop showing the check.
Turns the number into an action, not a dashboard nobody looks at.
Dispute or regret complaints, by prediction-accuracy cohort
12 / 100 6 / 100 0 2 / 100 Calibrated by listing 5 11 / 100 Stuck near chance
The lagging number, disputes weeks later, confirms exactly what the early guessing check already predicted well before it happened.

The recap, one line per letter: link is fewer disputes and less seller regret, early signal is the prediction-accuracy trend across the first ten listings, abuse is a seller gaming the check by always guessing their own number, and decision is a targeted re-explainer only for sellers still stuck at chance.

And if you want to be sure it really works, try it somewhere elseSame four letters, a code-review tool instead of a marketplace. Different domain, and the mental model being checked is about risk, not price.

Mergewise is a code-review tool that scores whether a pull request looks safe to auto-merge. Foster Amankwah is a backend engineer whose team turned on the auto-merge suggestion last quarter, and the real risk is a developer treating a high safety score as a certification instead of a probabilistic estimate.

Mapped onto LEAD: link is fewer bad auto-merges reaching production, the real cost the whole feature exists to avoid. Early signal is asking developers, right after onboarding, to predict whether Mergewise will flag a sample pull request as safe or risky before revealing the score, tracked across their first ten reviews. Abuse is a developer who always guesses "safe" since most pull requests genuinely are, looking accurate without ever engaging with the cases that actually diverge from that default. Decision is showing a real near-miss example to anyone still guessing "safe" every time by their tenth review, since that's the pattern most likely to miss a genuine risk later.

Hand sketched decision tree titled Mergewise, checking the mental model. Root: does the dev predict the merge risk score. Four branches: predicts well week two leads to stop nagging, guesses low every time leads to show a miss example, guesses high every time leads to show a safe example, never engages leads to short retry prompt.
Two of these four branches exist specifically to catch the developer whose guess never actually varies.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "measure whether they can predict it, not whether they click it," and stop.
Cost: there's no engineering time this quarter to build a live prediction-check screen. Say so, and start with a single post-onboarding quiz instead of a running check across ten listings.
The model gets better, for real: even as ValueMark's accuracy improves, sellers still need a mental model calibrated to when it's wrong, since a rarer miss on a better model is still a miss someone has to catch.

Where people run it wrong.
They measure acceptance or click-through rate and call it comprehension, when the two can look identical and mean completely different things.
They check comprehension once, right after onboarding, and never again, missing that the real test is whether it holds up over real use.
They build the check entirely from average-case examples, missing the edge cases the mental model actually needs to cover.

How to use it live. When someone asks how to measure a mental model, ask yourself first: what would this person have to predict, correctly, to prove they actually understand it? Measure that guess, not the click that follows it.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "how would you measure whether onboarding taught the right mental model"?
Tap to flip
ANSWER
LEAD: link, early signal, abuse, decision. It's a metric question, and the early signal step is the actual answer.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Bettina Voss, who sells her late grandmother's household items on Farrowgate most weekends.
3 · THE LINK
What real business outcome is this metric supposed to predict?
Tap to flip
ANSWER
Fewer disputes and fewer sellers who feel cheated after underselling something rare, both of which erode trust in the marketplace.
4 · THE EARLY SIGNAL
What is the actual leading-indicator metric?
Tap to flip
ANSWER
Prediction accuracy on a direction-and-size guess about ValueMark's price, made before the seller sees the real number, tracked across their first ten listings.
5 · THE REJECTED OPTION
What alternative metric was considered and rejected?
Tap to flip
ANSWER
Suggested-price acceptance rate. Rejected because a high acceptance rate looks identical whether a seller understands the tool or is simply clicking through.
6 · THE NUMBER
Fill in the blank: sellers stuck near chance filed disputes at a rate of ___ per 100 listings, versus 2 per 100 for calibrated sellers.
Tap to flip
ANSWER
11 per 100. The early prediction-accuracy gap showed up weeks before this lagging dispute number ever moved.
7 · THE ABUSE CASE
How could a seller game the prediction check without really understanding ValueMark?
Tap to flip
ANSWER
By always guessing "same as my own number." On cases where ValueMark agrees with naive intuition, that guess looks accurate without any real prediction happening.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what mental model is being checked there?
Tap to flip
ANSWER
Mergewise, a code-review tool. There, the check is whether a developer treats its merge-safety score as a probabilistic estimate rather than a certification.

Check yourself Score: 0 / 0

Multiple choice
1. Why doesn't this answer just use suggested-price acceptance rate as the metric?
  • A. Acceptance rate is too hard to measure.
  • B. A high acceptance rate looks the same whether a seller understands the tool or is just clicking through it.
  • C. Acceptance rate only applies to rare items.
  • D. ValueMark doesn't track acceptance rate at all.
Show hint
Look at the comparison diagram of the two mental models.
Show answer
B. Two sellers with the same acceptance rate can have completely different levels of real comprehension underneath it.
True or false
2. True or false: this answer recommends showing the prediction check on every single listing, forever.
  • True
  • False
Show hint
Look at "what I would leave alone."
Show answer
False. Once a seller's prediction accuracy is already high, the check stops, since more of it is just friction with nothing left to measure.
Fill in the blank
3. Fill in the blank: calibrated sellers' prediction accuracy climbed to about ___ percent by listing ten, up from around 52 percent at listing one.
Show hint
Look at the line chart of prediction accuracy by listing number.
Show answer
81 percent. Meanwhile sellers who never engaged stayed flat around 51 percent the entire time, essentially a coin flip.
Short answer, where it wouldn't matter
4. Name a seller for whom this prediction-check metric stops being useful, even though the check itself still exists.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A seller whose accuracy is already high by listing three. Continuing to check them adds friction without measuring anything new.
Short answer, name the abuse case
5. What's one way this metric could be gamed, and how does the design guard against it?
Show hint
Look at the "abuse" step in the LEAD recap.
Show answer
Model answer: A seller could always guess "same as my own number," looking accurate on average cases. The guard is weighting the check items toward cases where ValueMark actually diverges from naive intuition.
Short answer, apply it yourself
6. Pick a product you use yourself. What one prediction could it ask you to make, before showing you its answer, to test whether you actually understand how it thinks?
Show hint
Think about a recommendation feed, a spell-checker, or a maps app's time estimate.
Show answer
Model answer: A navigation app could ask you to guess which of two routes it will pick as faster before showing its answer. Getting better at that guess over time would mean you actually understand what it optimizes for.
Before you close the answer
Why this works
Tests whether you can find a metric that measures belief, not just behavior, and whether you understand why a behavioral proxy like acceptance rate can hide a completely broken mental model underneath a healthy-looking number.
Follow-up traps
"Isn't asking sellers to guess just extra friction with no real value?" Response: it's ten seconds, once per listing, only until accuracy is high, and it catches exactly the gap acceptance rate can't see.

"What if a seller just guesses randomly to get through the check faster?" Response: random guessing shows up as accuracy stuck near 50%, which is itself the signal that the mental model hasn't formed, so it doesn't break the metric, it becomes the metric.
If pressed
Farrowgate's real check items are drawn from a held-out set specifically chosen because ValueMark's price diverges from median naive-guess data by at least 15%, so a seller can't pass by defaulting to their own instinct.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more