InterviewAdvancedDesigning for Uncertainty & Trust / Trust, transparency and explainability in UX / #22

Design the transparency layer for a product I describe.

SPARK the product is TraceLine, an AI assistant that tells telecom field technicians the likely cause of a network fault

Here is the product: Harrow Telecom built TraceLine. A field technician stands at a broken piece of network gear, TraceLine reads the equipment's signal data, and it tells the technician what it thinks is wrong. Priya Vashisht has fixed roadside cabinets for six years and can usually guess a fault before her meter confirms it.

The direct answer
Every recommendation shows a three-part card, not a single score: the signal pattern TraceLine actually saw, its top guess with how sure it is, and the next most likely cause it rejected. Never ship a bare confidence number alone. A lone score invites a technician to stop checking, and that's the exact moment a wrong guess turns into a wrong repair.
Do this, in order
  1. Show the three-part card: signal, top cause with confidence, rejected alternative.Why: this is the one design decision the whole transparency layer depends on.
  2. Make the confidence band impossible to miss, not fine print under the top guess.Why: technicians need one glance to see when two causes are running close.
  3. Route any close-call recommendation into a mandatory signal check.Why: this is exactly the case that turns into a wrong part swap if nobody looks twice.
  4. Track the wrong-dispatch rate weekly, split by confidence band.Why: it's the leading number that catches over-trust creeping back before repair costs spike.
  5. Leave the technician's own signal check as the default step. No auto-apply mode.Why: the moment a person stops verifying is where the real cost hides.
  6. Keep the paper fault chart as a fallback for cabinets with no clean signal to read.Why: some rural sites genuinely have nothing for the card to explain.

How to answer this, stage by stage

You're not being graded on whether you remember every SPARK letter. You're graded on whether you can point at one design decision and defend it.

Stage 1
Scope the product to one person
Say it like this
"I'll design this for Priya, a field technician at Harrow Telecom, standing at a roadside cabinet with TraceLine open on her tablet."
Why this works
Turns "design a transparency layer" from an abstract phrase into one screen, one person, one decision.
Stage 2
Say your structure out loud
Say it like this
"I'll use SPARK. Situation, what she does today. Payoff, the habit I want to build. Anchor, the one design call. Risk, what breaks the day it's wrong. Keep out, what I won't build yet."
Why this works
Tells the interviewer you have a route through the answer before you start walking it.
Stage 3
Reframe the question
Say it like this
"A transparency layer isn't a disclaimer at the bottom of the screen. It's a design for the one day the model's top guess is wrong, and what the technician does next."
Why this works
Separates you from an answer that just adds a footnote and calls it done.
Stage 4
Give the one decision
Say it like this
"Every recommendation is a three-part card: the signal pattern, the top cause with a confidence band, and the next most likely cause. Never a bare score by itself."
Why this works
This matches the direct answer word for word. That match is what makes it defensible under follow-up.
Stage 5
Prove it with a failure
Say it like this
"In an early version that just showed a score, wrong-part swaps climbed from 6 percent to 21 percent over twelve weeks, because technicians stopped checking the signal themselves. The card fixes exactly that."
Why this works
A number the interviewer can picture beats a claim that the design is "safer."
Stage 6
Say what you'd leave alone
Say it like this
"I wouldn't touch the paper fault chart. Some rural cabinets don't emit clean signal at all, and no explanation card can explain data that was never there."
Why this works
Shows judgment instead of redesigning everything on principle.
Stage 7
Close on the one line
Say it like this
"Show the reasoning, not just a label. A score asks for total trust. A card lets trust be partial, which is the only kind that survives a wrong guess."
Why this works
Ends on the decision, not on a summary of the story you told to get there.

Let's learn

Harrow Telecom built TraceLine, an AI tool that reads the signal data coming off a broken piece of network gear and tells a field technician what it thinks caused the fault.

Before TraceLine, a technician like Priya Vashisht climbed up to a roadside cabinet, tested the line with a handheld meter, and paged through a forty-page paper fault chart to match the pattern to a known cause. That took about 47 minutes, on a good day.

Knowledge spark: what's a confidence band? A range around the model's top guess, not just one number. "70 to 85 percent sure" tells a technician there's real room for doubt. "78 percent" alone sounds more exact than it actually is.

An early version of TraceLine cut that 47 minutes down to 18. It read the signal and printed one line: the most likely cause, and a percentage next to it.

Hand sketched icon list titled Today, without TraceLine. Four items: a gauge icon labeled read the signal by hand, a document icon labeled flip through the paper chart, a person icon labeled call a senior tech if stuck, a box icon labeled guess, open panel, check.
This is what TraceLine had to replace: four slow, manual steps, each one a place a technician's own judgment used to sit.

The turn. The extra minutes TraceLine saved were never the real problem. The problem was what technicians stopped doing with those minutes: checking the signal themselves before opening a panel.

Diagnosis time down, wrong-dispatch rate up: the single-score version
50 25 0 47 Before 18 After Diagnosis time (minutes) 6% Before 21% After Wrong-dispatch rate
The single-score version really was faster. It was also nearly four times more likely to send a technician up the ladder with the wrong part.

At its worst: a technician would read "Corroded connector, 91%," swap the connector, and drive off without checking the actual line. Two weeks later, a second truck roll to the same cabinet, because the real fault was a cracked seal the first score never mentioned.

The decision I would take back We shipped TraceLine's first version showing a single line: top cause, plus a percentage. It made sense at launch, when the model's accuracy was new and every design instinct said "keep the screen simple." It stopped making sense once technicians learned to trust that single line more than their own meter.

What I would leave alone: the paper fault chart, kept as a backup for cabinets that emit no clean signal at all, usually older rural sites. No explanation card, however well designed, can explain data that was never there to begin with.

Hand sketched flow diagram titled The anchor, one explanation card. Four boxes in sequence: signal pattern, top cause with percent, alternative cause highlighted, suggested action.
The anchor is the third box. Showing the cause TraceLine ranked second is what gives a technician somewhere real to look when the first one is wrong.
A confidence score asks a technician to trust the whole card at once. A rejected alternative lets that trust be partial, which is the only kind that survives being wrong.

The lesson: a single number on a screen isn't a neutral fact. It's an invitation to stop checking, and a person who's good at their job will always take a fair invitation to work a little less.

Now here is the same thing as a story

The short version above is what you'd say in the room. Read this one for how the card actually got built.

Priya Vashisht has worked Harrow Telecom's northern routes for six years. Before TraceLine, she could often name a fault from the sound the cabinet fan made when she opened the panel, and she'd only pull out her meter to confirm it.

TraceLine's first months were good ones. Her average visit dropped from 47 minutes to 18. She started closing tickets before her lunch break instead of after it.

The habit thinned out in three beats. First, she stopped double checking cases the score called "high confidence." Then she stopped opening the paper chart at all, even for the odd ones. Then, without deciding to, she stopped touching her meter before swapping a part TraceLine named.

The trigger was small: a colleague on the southern crew, Jorik, mentioned over coffee that his second truck roll count had crept up all quarter. Nobody had connected it to TraceLine yet.

Hand sketched comparison diagram titled The day the top guess is wrong. Left panel, a gauge icon labeled Score only, caption trusted, wrong part swapped. Right panel, a document icon labeled Full card, caption checked, caught in time.
Same wrong top guess, two designs. One sends a technician up the ladder twice. One catches it on the ground.

Priya pulled her own numbers that week. Wrong-part swaps, cases where the part she'd replaced turned out not to be the actual fault, had climbed from about 6 percent of her visits to 21 percent over twelve weeks.

Hand sketched timeline titled Twelve weeks, trust becomes over-trust. Four milestones: rollout week 1, top guess trusted week 6, wrong dispatch spike highlighted week 10, anchor redesign week 14.
Nobody decided to stop checking on any single day. It happened between week one and week six, and nobody wrote it down.
Wrong-dispatch rate, week by week
25% 12% 0 week 10: 21% anchor ships week 16: 7% wk1 wk16
The rate kept climbing for four weeks after the redesign shipped, since cards already in a technician's queue still used the old design. It settled near 7 percent by week 16, close to where it started.

She raised it at the next regional call, not as a complaint about TraceLine's accuracy, since the model itself hadn't gotten any worse, but as a question about what the screen was actually asking her to do.

The redesigned card shipped in week fourteen: signal pattern, top cause with a confidence band, and the cause TraceLine ranked second. The team considered a stricter fix first, a mandatory second confirmation step on every single recommendation, and rejected it. It would have erased the 29 minutes TraceLine saved on every visit, which defeats the reason the tool exists.

Hand sketched labeled parts diagram titled What we left for later. Center box labeled Not day one, with four callouts: auto-apply fix, cross-tech compare, drift dashboard, auto-close ticket.
None of these four were wrong ideas. They were just riskier than the team needed to be on day one.

Same fault, same cabinet, replayed with the new card: Priya sees "Corroded connector, 68 to 79 percent" next to "Cracked seal, second most likely." The band is narrow enough to make her check the seal before she opens her toolbox. Four minutes added to her visit. Zero return trips.

I built the single-score version because it tested well in a demo and looked clean on a small screen. It took watching Priya's own wrong-dispatch count nearly quadruple, on a model that hadn't changed at all, to see that the card was never really about looking clean. It was about giving her a reason to keep doing the thing she was already good at.

SPARK, in one screenNot a story about a mistake. SPARK is what forces the design to survive the day it's wrong, before you ship it.

S
Situation. Her day, without you.
Priya climbs to a cabinet, tests the line by hand, checks a 40-page paper chart. 47 minutes, one cabinet.
Grounds the whole design in one real, unglamorous workflow.
P
Payoff. The habit you want built.
Not "trust the tool faster." Read the reasoning behind a recommendation before acting on it.
Names the behavior the product should build, not just the time it should save.
A
Anchor. The one design call.
Every recommendation is a three-part card: signal, top cause with confidence, rejected alternative.
The single decision everything else in the answer defends.
R
Risk. What breaks when it's wrong.
Wrong-dispatch rate climbed from 6% to 21% under the score-only design over 12 weeks.
Proves the anchor was built to survive the exact failure it names.
K
Keep out. What's not day one.
No auto-apply fix, no cross-technician comparison, no full drift dashboard. Not yet.
Shows judgment about scope instead of pitching every feature at once.
Hand sketched quadrant titled What ships day one, by risk and trust value. Axes: how risky to ship now from safe to risky, how much it changes trust from small to large. Confidence band and alternative cause sit top left, safe and high trust value. Drift dashboard and auto-apply fix sit bottom right, riskier and lower near-term value.
The two pieces of the card sit safely in the top left. Auto-apply and the full dashboard sit in the corner that earns the "not yet."

The recap, one line per letter: situation is Priya's 47-minute manual diagnosis, payoff is reading the reasoning instead of the label, anchor is the three-part card, risk is the wrong-dispatch spike the score-only version caused, and keep out is the auto-apply mode the team deliberately didn't ship.

And if you want to be sure it really works, try it somewhere elseSame five letters, a disaster-relief nonprofit instead of a telecom crew. A different flip family shows up here too: this one is about a delegation risk, not over-trust.

Bridge Relief Network runs IntakeLine, an AI tool that reads a family's shelter-intake answers and flags urgent needs for the coordinator on shift. Esteban Duarte manages intake at one of their busiest sites, and used to interview every arriving family himself before IntakeLine existed.

Mapped onto SPARK: situation is Esteban assessing each family by hand on paper forms, asking the same trauma-sensitive questions over and over across shifts. Payoff is coordinators trusting the shared record enough to stop re-asking those questions at every handoff. Anchor is that every auto-suggested flag, like "high priority: medical need," shows the exact answers that triggered it, with the underlying fields editable, not just an accept-or-reject button. Risk is the day IntakeLine misses something because a family answered indirectly; the anchor design lets a coordinator see the raw answers behind the flag and catch the gap immediately, instead of trusting a flag that quietly said nothing was wrong. Keep out is any automatic bed or resource assignment straight from the flag. A person still makes that call.

Hand sketched decision tree titled IntakeLine flags a need, what a coordinator does next. Root: IntakeLine flags a priority. Three branches: coordinator confirms it leads to keep the tag, coordinator disagrees leads to edit note why, coordinator unsure leads to escalate to shift lead.
Three branches, and none of them let a flag become a decision on its own. A person always closes the loop.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "show the reasoning, not the label, so trust can be partial," and stop.
Cost: there's no budget this quarter for a full redesign. Say so honestly, and ship the rejected alternative first; it's the cheapest single addition and it catches most of the close calls on its own.
The model gets better, for real: if TraceLine's overall accuracy improves next quarter, that's not a reason to remove the card. A better model still has a worst 5 percent of cases, and those are exactly where a technician needs the reasoning most.

Where people run it wrong.
They treat "transparency" as a disclaimer added after the design is done, not a decision made before it.
They assume more information always helps, and bury the one useful signal, the rejected alternative, under ten fields nobody reads.
They design for the demo case, where the model is right, instead of the case that actually matters: the one where it's wrong.

How to use it live. When someone asks you to design a transparency layer, ask yourself one question first: what does this person do the day the model's top guess is wrong? Design backward from that answer, not forward from what looks clean on a screen.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "design the transparency layer for a product I describe"?
Tap to flip
ANSWER
SPARK: situation, payoff, anchor, risk, keep out. The anchor step is the answer to the question itself.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Priya Vashisht, a field technician who has fixed roadside cabinets for Harrow Telecom for six years.
3 · THE PAYOFF
What habit is the design actually trying to build?
Tap to flip
ANSWER
Reading the reasoning behind a recommendation before acting on it, instead of trusting a single label.
4 · THE ANCHOR
What's the one concrete design decision in this answer?
Tap to flip
ANSWER
A three-part card on every recommendation: the signal pattern, the top cause with a confidence band, and the next most likely cause.
5 · THE RISK
What broke under the single-score version?
Tap to flip
ANSWER
Wrong-dispatch rate climbed from 6% to 21% over 12 weeks as technicians stopped checking the signal themselves.
6 · THE NUMBER
Fill in the blank: TraceLine cut diagnosis time from 47 minutes to ___ minutes.
Tap to flip
ANSWER
18 minutes. That speed is exactly what made the single-score shortcut so tempting to lean on.
7 · THE REPLAY
Same wrong top guess, redesigned card. What changes?
Tap to flip
ANSWER
Priya sees a narrow confidence band and a named alternative, checks the seal before opening her toolbox, adds 4 minutes, and avoids a second truck roll entirely.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's different about the anchor?
Tap to flip
ANSWER
Bridge Relief Network's IntakeLine. There, the anchor shows the raw answers behind a flag so a coordinator can catch what the flag missed, not just a rejected alternative cause.

Check yourself Score: 0 / 0

Multiple choice
1. Why does the anchor design show a rejected alternative cause instead of only the top guess?
  • A. Because the model can't produce just one cause.
  • B. Because it gives the technician something concrete to check when the top guess turns out wrong.
  • C. Because regulators require at least two causes to be listed.
  • D. Because it makes the screen look more detailed to new hires.
Show hint
Look at "the day it's wrong" comparison diagram.
Show answer
B. The alternative cause is the one thing on the card built specifically to survive the model being wrong.
True or false
2. True or false: this answer recommends adding a mandatory second-person check on every single TraceLine recommendation.
  • True
  • False
Show hint
Look at the rejected alternative in the story, the "stricter fix" the team considered.
Show answer
False. That option was considered and rejected, since it would have erased the 29 minutes TraceLine saves per visit. The card, not a mandatory second check, is the actual anchor.
Fill in the blank
3. Fill in the blank: over twelve weeks, the wrong-dispatch rate climbed from 6 percent to ___ percent.
Show hint
Look at the grouped bar chart in Section 1.
Show answer
21 percent. Nearly four times higher, on a model whose own accuracy hadn't changed at all.
Short answer, name the decision
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Shipping a single score with no reasoning shown. It made sense at launch, when keeping the screen simple felt like the safer call and the model's accuracy was still new.
Short answer, where it wouldn't matter
5. Name a place in this same product where a fancier transparency layer wouldn't help at all.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A cabinet with no clean signal to read at all. No card can explain data that was never captured.
Short answer, apply it yourself
6. Pick an AI product you use yourself. What's one place it gives you a bare label or score with no reasoning behind it, and what would you add?
Show hint
Think of a spam filter, a spell checker, or a recommendation you were never told the reason for.
Show answer
Model answer: A spam folder that just says "spam" with no reason shown. Even one line, "flagged: sender not in your contacts," would let you catch the rare miss yourself.
Before you close the answer
Why this works
Tests whether you design a transparency feature around the day the model is right, which is easy, or the day it's wrong, which is the only day that actually matters.
Follow-up traps
"Isn't showing a rejected alternative just going to confuse technicians?" Response: only if it's buried; keeping it to one line, ranked clearly below the top cause, is what keeps it useful instead of noisy.

"What if the model is right almost every time, doesn't this slow everyone down for nothing?" Response: the four extra minutes only really get spent when the confidence band is already narrow, which is precisely the case where being slow once beats a second truck roll.
If pressed
Harrow's real rollout only forces the mandatory signal check when the top two causes sit within 15 points of confidence of each other, so the added friction lands on the genuinely close calls, not on every visit.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more