Artifact critiqueIntermediateQuality, Cost & Token Economics / Quality metrics: accuracy vs usefulness vs trust / #9

Design a survey instrument to measure perceived reliability.

A survey instrument built around a single specific request the person can actually remember, checked against what the system logs say really happened.

The direct answer
Anchor every question to one specific, recent, real request the respondent made, never a general "how reliable is it" scale. Ask what they asked for, ask if it did that exact thing, and match their answer against the system log for that same request. A survey that asks about "overall" reliability measures the respondent's mood that day, not the product. A survey anchored to one real event measures the product, and you can check it.
Do this, in order
  1. Anchor every question to one specific, recent, real interaction, never a general reliability scale.Why: a vague scale question has no fixed thing to remember, so people answer with a feeling instead of a fact.
  2. Keep the recall window short, a few days, not a quarter.Why: memory of one small request fades fast; a three-month lookback turns into a guess dressed up as a fact.
  3. Build a way to match the response back to the actual log of that one request.Why: this is the only way to tell a real failure from a person misremembering, or from the model quietly doing the wrong thing while reporting success.
  4. Split results by what kind of request it was, not one blended score.Why: a scene-and-routine failure and a single-light failure are different problems, and a blended score hides which one is actually breaking.
  5. Set a real threshold for when a category's self-reported failure rate counts as a signal, not noise.Why: without a bar, every small dip in the number turns into a fire drill, and real drift gets lost in that noise.
  6. Leave the general satisfaction question in place for things that really are about overall feeling, like price or design.Why: not every question needs an anchor; forcing one onto a genuine opinion question just makes the survey longer for no reason.

How to answer this, live

Nobody is grading whether you can name "ask about a recent interaction" as a phrase. They're grading whether you can show why a vague reliability question is actually broken, out loud, in front of them. Seven moves get you there.

1
Name the product before you design anything
Say it like this
"Let's make this concrete. Briarwood Home makes a voice hub. You ask it to turn on a lamp, set the thermostat, or run a whole scene like 'movie night,' and it just does it. I'm going to design the survey for that."
Why this works
A survey design pitched at "an AI product" in general is a survey nobody could actually ship. Naming the real thing forces every later choice to be concrete.
2
Say what's wrong with the obvious version first
Say it like this
"The obvious move is one question: 'Overall, how reliable is your Briarwood Hub, one to five?' I'd reject that on purpose. It has no fixed thing to remember, so people answer with how their week is going, not with a fact about the product."
Why this works
Naming and rejecting the easy answer shows judgment, not just a different opinion. This is the alternative you considered and turned down.
3
State your structure out loud
Say it like this
"I'd run this through SPARK. Ground it in what the person does today without a good survey, decide the habit I want the number to build, pick the one design decision everything hangs on, plan for the day that decision is wrong, and say what I'm deliberately not trying to catch this way."
Why this works
Two seconds of structure tells the interviewer you have a repeatable method for building a measurement tool, not just a clever question idea.
4
Give the anchor, the actual answer
Say it like this
"Ask them about the last time, in the past three days, they asked the Hub to do something. What did you ask for? Did it do that exact thing? Then, with their okay, I match that answer against the log for that same request."
Why this works
This is the direct answer, said the way you'd actually say it out loud. Everything after this is why it's the right design, not a different one.
5
Say what breaks the first time you're wrong
Say it like this
"If I just anchor the question but never check it against the log, I've only fixed the recall problem, not the trust problem. The Hub can log 'success' on a request it actually botched, and the person's own memory of what happened won't catch that on its own."
Why this works
Naming the specific way your own design fails is what separates a real decision from a slogan. It shows you've already thought past the happy path.
6
Say what you're deliberately not asking
Say it like this
"I wouldn't ask whether they've stopped using voice for scenes and started tapping the app instead. That's a behavior, and behavior belongs in product analytics. A survey can tell you how one request felt. It can't tell you what someone quietly stopped doing afterward."
Why this works
Naming a boundary on purpose shows you understand what a self-report instrument is actually for, instead of trying to make it do everything.
7
Close on the threshold, and what it would have caught
Say it like this
"The test I hold myself to: I set a real bar, something like 15 percent self-reported failure in a category, checked against at least 40 anchored, log-matched responses in two weeks, before it counts as a real signal. At Briarwood, that bar would have caught the scene-command regression in week three. The old survey didn't catch it until week ten, from support tickets."
Why this works
Closing on a real number, not a feeling, is what turns "anchor it to a real event" from a nice idea into a design you could actually defend in a review.
If you remember one thing A reliability survey is only honest if you can point to the one real thing it's asking about. No anchor, no honest number.

The plain version

What happens when a survey says your product is fine, and the product is not fine?

Briarwood Home makes a voice hub. It sits on a shelf or a counter. You tell it to turn on a lamp, warm up the living room, or run a whole scene, like "movie night," which is really six small commands stacked into one word.

For three years, Briarwood asked its customers one question every quarter, by email: "Overall, how reliable is your Briarwood Hub, from one to five?" About 9,000 people answered each time, out of roughly 180,000 households. The score sat near 4.2. Steady. Nobody worried about it.

The score wasn't broken. It just was never asking about the thing that was actually going wrong.
Knowledge spark: what is a silent wrong action? Sometimes a voice assistant hears you fine, picks the wrong device or the wrong scene anyway, does that wrong thing, and still tells you "Done." Nothing errors. Nothing asks you to check. The log says success. You just end up with the wrong light on.

At its worst, a survey like this can sit calm for months while something real breaks underneath it, because a general "how reliable is it" question has no fixed thing anyone is picturing when they answer. Someone who had a great morning and someone whose living room lamp just turned on by itself can both circle a 4.

The decision that mattered When Briarwood built this survey, two years earlier, nobody linked a response back to the actual request it was about. Legal wanted the survey fully separate from any activity log, to keep it simple and clean. It made sense with a small user base and no team built yet to match survey answers to logs. Nobody revisited it once that team existed.

What I would leave alone: a general "how do you feel about the Hub" question is genuinely fine for things like price or the way the device looks on a shelf. Those really are about overall feeling, not one fact you could check. The mistake was using that same soft question for something as checkable as "did it do what I asked."

The lesson: a survey question is only as good as the memory it asks for. Ask about a feeling in general and you get today's mood. Ask about one real thing that happened this week and you get a fact you can go check.

The same thing, as a story

The short version sits above. Read on for the quarterly review where a support director asked Roksana a question the survey couldn't answer.

Roksana Anholt built Briarwood's quarterly reliability survey herself, back when the company had one product line and about 40,000 households. For most of two years it did exactly what it was supposed to do. It sat near 4.2, and when something genuinely broke company-wide, it dipped, just a little, just enough to matter.

Then Briarwood shipped a firmware update that changed how the Hub resolved a request when a scene touched more than one device with a similar name, like two lamps in the same room. The update was meant to make scenes faster to build. It also made the Hub guess wrong more often about which lamp a scene meant, and when it guessed wrong, it still said "Done," because from the Hub's own point of view, it had completed a scene. It just wasn't the scene the person meant.

Nobody noticed for a while, because nothing about the number Roksana watched gave her a reason to look. Support tickets mentioning a scene "doing the wrong thing" climbed from about 40 a week to around 130 a week over ten weeks. The quarterly survey, over that same stretch, moved from 4.3 to 4.2 to 4.2 to 4.1. Inside the range it always sat in.

The trigger wasn't a crisis. It was one sentence in a quarterly review. The VP of Support looked at the reliability slide, then at her own inbox, and said, "If this has been flat at 4.2 for a year, why is my scene-complaint queue three times bigger than it was in the spring?"

Roksana didn't have an answer. She pulled the raw text responses from the last three surveys to look for a pattern by hand. A good chunk of respondents had written some version of "I don't really use scenes, not sure what you mean." The number wasn't hiding a scene problem. It had never been asking about scenes in the first place.

The score wasn't measuring the Hub. It was measuring whether people were having an alright week.

Worse, she couldn't check a single complaint-shaped comment against what actually happened, because the survey had been built two years earlier with no way to connect one person's answer to their own activity log. That decision made sense at launch, when the team was four people and there was no golden-set pipeline to feed anyway. Nobody had come back to change it once the company had both.

Hand sketched diagram titled today no anchor guessing why the score will not move. Center figure labeled Roksana before the redesign. Four callouts around her: one vague scale question, no link to the real event, same email each quarter, score reads fine anyway.
What the old instrument actually was: one soft question, asked the same way every quarter, with no way to tie any single answer back to a real request.

So Roksana rebuilt the survey around one decision. Instead of asking how reliable the Hub is overall, the new version asks: "Think about the last time, in the past three days, you asked your Hub to do something. What did you ask for?" Then, about that one specific thing: "Did it do that exact thing?" And, with a clear opt-in, "Can we check this against your Hub's activity log for just this one request?"

Hand sketched diagram titled the anchor one card one instruction. Center icon a document labeled the anchored question. Four callouts: your last real request past 3 days, what did you ask for, did it do that exact thing, optional match to the log.
The whole design hangs on this one card. Every other question in the survey is optional. This one is not.

The next time a similar regression shipped, seven months later, Roksana reran the comparison on purpose, as a test. Under the old survey, the flat quarterly score gave no warning until support tickets forced the issue, around week ten. Under the anchored, log-matched survey, the self-reported failure rate among people whose anchored request was a scene command crossed 15 percent by week three, on a sample of about 60 responses, well past the bar she'd set for a real signal. Of those, log matching confirmed roughly 90 percent were genuine mismatches, not people misremembering. Seven weeks of exposure became three.

Blended survey score vs. anchored, log-checked scene failure rate, week 1 to week 10
40% 0% crosses 15% bar Wk 1 Wk 4 Wk 8 Wk 10
Blended survey score (scaled to %)Anchored scene-category failure rate
The blended score barely moves the whole time, staying near its usual range. The anchored, log-checked line for scene commands climbs past the 15 percent bar by week 3, seven weeks before the old survey ever reflected a problem.

What the anchor design deliberately does not try to catch: whether people quietly gave up on saying "movie night" out loud and started tapping the scene button in the app instead. That's a real trust signal, and it matters, but it lives in Briarwood's own usage logs, not in anything a person would think to volunteer on a survey about one recent request.

SPARK, in one screen

Not a diagnosis of one bad quarter. SPARK run on the actual design of the instrument, using a real regression as the way to check whether the design would have worked.

S
Situation. How the job gets done today, without a good instrument.
Roksana defended a single flat number every quarter with no way to explain what, specifically, it did or did not cover. She could not tell a director which command category was fine and which wasn't, because the survey was never built to say.
This grounds the whole design in a real moment: a PM standing in a review with a number she cannot actually defend.
P
Payoff. The habit the redesign should build.
Stop reaching for one soft number in a review. Start reaching for a specific category and a real rate: "scene commands are failing for 19 percent of people who tried one in the last two weeks," not "reliability is 4.2."
The habit is the actual product here. A better number that nobody changes their behavior around isn't worth building.
A
Anchor. The one decision everything else hangs on.
Every question ties to one specific, recent, real request the respondent made, never a general scale. "Think of your last request in the past three days" replaces "how reliable is it overall." This is the actual answer to the question asked.
Everything downstream, the categories, the threshold, the log match, only works because this one decision gives every later question something fixed to point at.
R
Risk. What breaks the first time the design is wrong.
An anchored question alone still trusts the person's memory completely. The failure mode: the Hub logs "success" on a request it silently got wrong, so a purely self-reported anchor can't tell a real failure from someone half-remembering. Guardrail: pair every anchored response with an opt-in match against the actual command log for that same request, and treat a logged-success-but-reported-failure mismatch as its own escalation category, added to the set of real examples the model team retests against for that intent, not averaged away.
This is the strongest move in the whole design. It's what turns "ask about a specific event" from a nicer survey into something you can actually trust.
K
Keep out. What the survey deliberately does not try to catch.
Behavioral trust signals: whether someone quietly stopped using voice for a command type and switched to tapping the app instead, or how often they retry a request before giving up. Those live in product usage logs, which already track them, not in a self-report instrument nobody would think to fill out about a thing they didn't do.
Naming the boundary on purpose is what keeps the survey from trying to be three instruments at once and doing all three badly.
Hand sketched comparison diagram titled same regression two instruments. Left panel labeled old design, caption score sits near 4.2 for 10 weeks straight. Right panel labeled anchored design, caption scene failures cross 15 percent by week 3 log checked.
Same regression, run through two instruments. One instrument stays quiet the whole time. The other one rings early enough to matter.

Three things worth saying plainly, since this is where the real judgment sits. The alternative rejected was the industry-standard single reliability scale, the exact question Briarwood had used for three years: it lost because it cannot be traced to a command category, it stays flat regardless of what's actually breaking underneath it, and it is genuinely vulnerable to the respondent's mood that day rather than any fact about the product. The AI-specific failure named is silent wrong execution, a form of the Hub confidently doing the wrong thing and reporting it as done, which no ordinary satisfaction survey design would ever think to check for, because the guardrail against it (log matching) only makes sense once you already know your own model can be confidently wrong. The trade-off accepted on purpose: an anchored, three-day recall window and a short, tightly scoped survey trades some of the depth a longer, more reflective quarterly survey could offer, for a much higher chance that what people report actually happened the way they say it did. And the threshold that turns a number into a real signal, not morning noise: a command category only counts as a live problem once its anchored, log-matched self-reported failure rate crosses roughly 15 percent within a rolling two-week window, on at least 40 anchored responses for that category, not on one bad batch of five.

Run it again, somewhere with no smart speaker in sight

Same five letters, a city's 311 phone line instead of a smart hub, and the anchor idea holds up with a completely different kind of AI system underneath it.

The City of Aldercreek runs an AI phone line for 311. Residents call to report a pothole or schedule a bulk trash pickup, and the assistant handles most calls without a human ever picking up. Colette Deschamps analyzes call data for the city's public works department.

S, situation. Today, without a good instrument, Aldercreek relies on the automated post-call rating: press a number, one to five, right after you hang up. It has averaged 4.4 for over a year. Colette has no way to tell whether that number reflects the assistant actually doing the right thing, or just sounding polite and answering fast.
P, payoff. Stop citing "4.4 out of 5, satisfied" in a public works meeting. Start citing whether the specific thing a caller asked for actually got logged the way they asked for it.
A, anchor. A short follow-up text two days after the call: "What did you call about?" then "Did what happened match what you were told would happen?" matched, with consent, against the actual work order created from that call.
R, risk. The old post-call rating mostly measures how pleasant the voice sounded and how short the hold was, not whether the request landed correctly. The AI-specific failure here: the assistant mishears a cross-street name and logs a bulk pickup request under the wrong address, then confirms the pickup out loud, so the caller hangs up thinking it worked. Guardrail: log-matching catches the address mismatch even when the caller never noticed anything wrong.
K, keep out. How many days it actually takes public works to clear a logged pothole is a real operational number, already tracked in the work-order system. It doesn't belong in a phone survey about whether the call itself went the way the caller expected.

Bulk-pickup requests, address-mismatch rate: old post-call rating vs. anchored, log-matched follow-up
25% 0% 1% 22% Old post-call rating Anchored, log-matched
Implied problem rate from a flat 4.4 scoreReal address-mismatch rate, first two weeks
A steady 4.4 rating implies almost nothing is wrong. The anchored follow-up, checked against the actual work order, found 22 percent of bulk-pickup calls had been logged under the wrong address, invisible to the star rating the whole time.

Swap the trigger and it still runs.
Speed: an interviewer gives you ninety seconds. Skip straight to the anchor: "ask about one real request, in the last few days, and check it against the log." Everything else is why that one line is correct.
Cost: the log-matching pipeline is eight weeks out. Don't drop the anchor while you wait, keep asking the anchored question and read the free text by hand for outright contradictions until the pipeline ships.
The model got better, for real: say the intent-recognition model's own offline accuracy genuinely improved that quarter. That still isn't proof every category is fine. A model that is more accurate on average can still be badly wrong on one category it wasn't tuned for, and neither the old scale nor an improved accuracy score would show you that, only the anchored, log-matched category breakdown would.

Where people get this wrong.
They add the anchored question, then still lead every review with the old blended score, because that's the one number leadership already knows how to read.
They treat a stable anchored score as proof of nothing wrong, instead of checking whether the category it's measuring is even getting enough real traffic to say anything yet.
They respond to one bad category by adding a fifth or sixth survey question instead of asking whether the real fix is a guardrail in the product, not another question on the form.

How to use it live. Open with the anchor, not the framework name: "I'd ask about one real recent request, not a general feeling, and check it against the log." That gives you a concrete decision on the table before the interviewer can steer you toward a longer, softer list of survey questions.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
Which framework fits a "design a survey instrument" question like this one?
Tap to flip
ANSWER
SPARK: ground the design in today's real workflow, name the habit it should build, pick the one anchor decision, plan for the day it's wrong, and say what you deliberately leave for product analytics instead.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Roksana Anholt, an insights PM at Briarwood Home, a voice hub for controlling lights, thermostats, and scenes at home.
3 · THE OLD HABIT
What habit had Roksana leaned on that stopped being enough?
Tap to flip
ANSWER
Trusting one flat quarterly reliability score without ever being able to explain which specific command category it did, or didn't, actually cover.
4 · THE ANCHOR
What's the one design decision every question in this survey hangs on?
Tap to flip
ANSWER
Anchor every question to one specific, recent, real request the respondent made, like "your last request in the past three days," instead of a general "how reliable is it" scale.
5 · THE OLD DECISION
What old decision does this answer take back, and why did it make sense when it was made?
Tap to flip
ANSWER
Building the original survey with no link back to the respondent's own activity log. It made sense at launch, with a tiny team and no golden-set pipeline to feed. Nobody revisited it once both existed.
6 · THE NUMBER
Fill in the blank: under the anchored, log-matched survey, the scene-command failure rate crossed the 15 percent bar by week ___, against week ___ under the old survey.
Tap to flip
ANSWER
Week 3; week 10. Seven weeks of exposure became three, and the old survey only caught the problem because of a support ticket spike, not because of anything it measured directly.
7 · THE REPLAY
Same firmware regression, run again seven months later with the redesigned survey. What changes?
Tap to flip
ANSWER
The anchored, log-matched scene-category failure rate crosses its 15 percent bar by week 3 on about 60 responses, with roughly 90 percent of reported failures confirmed real by the log match, instead of surfacing through support tickets seven weeks later.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the parallel?
Tap to flip
ANSWER
The City of Aldercreek's 311 phone line, analyzed by Colette Deschamps. Same SPARK anchor idea: ask about one real recent call and check it against the work-order log, which caught a 22 percent address-mismatch rate the old post-call star rating never showed.

Check yourself Score: 0 / 0

Multiple choice
1. Why does anchoring every question to one specific, recent, real request work better than asking "how reliable is it overall"?
  • A. It makes the survey shorter, which is the only thing that matters for response rate.
  • B. It gives the respondent one fixed, real thing to remember, instead of inviting them to answer with their general mood.
  • C. It lets you skip building any way to check the response against real data.
  • D. It is required by most survey software by default.
Show hint
Think about what a person is actually picturing in their head when they answer a vague "overall" question versus a specific one.
Show answer
B. A vague scale question has no fixed event behind it, so the answer tracks how the respondent feels in general that day. An anchored question forces them to recall one real thing, which is what makes the answer checkable.
True or false
2. True or false: switching the old reliability question from a 1-to-5 scale to a wider 1-to-10 scale would fix the real problem with it.
  • True
  • False
Show hint
Ask what a wider scale actually changes about what the respondent is picturing when they answer.
Show answer
False. The problem was never the number of points on the scale. It was that the question had no anchor to a real event, so a wider scale still measures the respondent's mood, just with more decimal places.
Fill in the blank
3. In the design's threshold rule, a command category counts as a real signal once its anchored, log-matched self-reported failure rate crosses about ___ percent within a rolling two-week window, on at least ___ anchored responses for that category.
Show hint
Look at the closing paragraph of "SPARK, in one screen," right after the rejected alternative and the AI-specific failure mode.
Show answer
15 percent; 40 responses. Without a real bar like this, a single bad week in a small sample would trigger a false alarm every time.
Short answer, name the rejected alternative
4. The single "how reliable is it overall, one to five" question was Briarwood's original design, and it's a real, standard-looking metric. Why was it rejected as the anchor for this redesign?
Show hint
Look at what that question can and can't be traced back to, in the framework recap section.
Show answer
Model answer: It can't be traced to any specific command category, so it can't say which kind of request is actually breaking. It stays flat regardless of what's failing underneath it, because it has no fixed event to anchor the answer to, so it mostly reflects the respondent's mood that day rather than a fact about the product.
Short answer, apply it yourself
5. Pick an AI product you use yourself. What would the anchored question look like for it, and what would you check the answer against?
Show hint
Ask what the smallest recent request or interaction is that you could actually remember clearly, and whether there's a log anywhere that could confirm what really happened.
Show answer
Model answer: A grocery app with an AI substitution feature. Anchor: "Think of your last delivery in the past week where an item was out of stock and swapped. What did you order, and what did you get instead?" Check it against the order's own substitution log, matched by order ID, to confirm whether the swap the customer remembers matches what the system actually shipped.
Multiple choice
6. What does this survey design deliberately leave out, on purpose, rather than try to capture through self-report?
  • A. Which command category a request belonged to.
  • B. Whether the respondent's recent request succeeded or failed.
  • C. Whether someone quietly stopped using voice commands for a task and switched to the app instead.
  • D. The date of the respondent's last request.
Show hint
Look at the "K, keep out" step in the framework recap. It's a behavior, not something a person would think to report.
Show answer
C. That's a behavioral trust signal, already tracked in product usage logs. A self-report survey can only tell you how one remembered request felt, not what someone quietly stopped doing afterward.
Before you close the answer
Why this works
Tests whether you can design a self-report instrument that survives being checked against reality, not just write a plausible-sounding survey question. Most candidates stop at "ask about a specific recent interaction" and never say how they'd know the answer was true.
Follow-up traps
"What if people just don't remember their last request accurately, even within three days?" Response: that's exactly why the log match exists. The anchor narrows what they're trying to recall to one small, recent thing, and the log match catches the cases where memory still slips.

"Isn't cross-referencing survey answers against activity logs a privacy problem?" Response: it's opt-in per response, tied to that one specific request only, not a standing link between someone's identity and their full history. Someone can answer the anchored question honestly and decline the log match, and you still get more than the old blended score ever gave you.
If pressed
The actual gating rule used at Briarwood: a category only escalates to the model team's golden set once its mismatch rate, logged success but self-reported failure, holds above 10 percent of that category's anchored responses for two straight weekly pulls, not one.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more