Design the transparency layer for a product I describe.
Here is the product: Harrow Telecom built TraceLine. A field technician stands at a broken piece of network gear, TraceLine reads the equipment's signal data, and it tells the technician what it thinks is wrong. Priya Vashisht has fixed roadside cabinets for six years and can usually guess a fault before her meter confirms it.
- Show the three-part card: signal, top cause with confidence, rejected alternative.Why: this is the one design decision the whole transparency layer depends on.
- Make the confidence band impossible to miss, not fine print under the top guess.Why: technicians need one glance to see when two causes are running close.
- Route any close-call recommendation into a mandatory signal check.Why: this is exactly the case that turns into a wrong part swap if nobody looks twice.
- Track the wrong-dispatch rate weekly, split by confidence band.Why: it's the leading number that catches over-trust creeping back before repair costs spike.
- Leave the technician's own signal check as the default step. No auto-apply mode.Why: the moment a person stops verifying is where the real cost hides.
- Keep the paper fault chart as a fallback for cabinets with no clean signal to read.Why: some rural sites genuinely have nothing for the card to explain.
How to answer this, stage by stage
You're not being graded on whether you remember every SPARK letter. You're graded on whether you can point at one design decision and defend it.
Let's learn
Harrow Telecom built TraceLine, an AI tool that reads the signal data coming off a broken piece of network gear and tells a field technician what it thinks caused the fault.
Before TraceLine, a technician like Priya Vashisht climbed up to a roadside cabinet, tested the line with a handheld meter, and paged through a forty-page paper fault chart to match the pattern to a known cause. That took about 47 minutes, on a good day.
An early version of TraceLine cut that 47 minutes down to 18. It read the signal and printed one line: the most likely cause, and a percentage next to it.
The turn. The extra minutes TraceLine saved were never the real problem. The problem was what technicians stopped doing with those minutes: checking the signal themselves before opening a panel.
At its worst: a technician would read "Corroded connector, 91%," swap the connector, and drive off without checking the actual line. Two weeks later, a second truck roll to the same cabinet, because the real fault was a cracked seal the first score never mentioned.
What I would leave alone: the paper fault chart, kept as a backup for cabinets that emit no clean signal at all, usually older rural sites. No explanation card, however well designed, can explain data that was never there to begin with.
The lesson: a single number on a screen isn't a neutral fact. It's an invitation to stop checking, and a person who's good at their job will always take a fair invitation to work a little less.
Now here is the same thing as a story
The short version above is what you'd say in the room. Read this one for how the card actually got built.
Priya Vashisht has worked Harrow Telecom's northern routes for six years. Before TraceLine, she could often name a fault from the sound the cabinet fan made when she opened the panel, and she'd only pull out her meter to confirm it.
TraceLine's first months were good ones. Her average visit dropped from 47 minutes to 18. She started closing tickets before her lunch break instead of after it.
The habit thinned out in three beats. First, she stopped double checking cases the score called "high confidence." Then she stopped opening the paper chart at all, even for the odd ones. Then, without deciding to, she stopped touching her meter before swapping a part TraceLine named.
The trigger was small: a colleague on the southern crew, Jorik, mentioned over coffee that his second truck roll count had crept up all quarter. Nobody had connected it to TraceLine yet.
Priya pulled her own numbers that week. Wrong-part swaps, cases where the part she'd replaced turned out not to be the actual fault, had climbed from about 6 percent of her visits to 21 percent over twelve weeks.
She raised it at the next regional call, not as a complaint about TraceLine's accuracy, since the model itself hadn't gotten any worse, but as a question about what the screen was actually asking her to do.
The redesigned card shipped in week fourteen: signal pattern, top cause with a confidence band, and the cause TraceLine ranked second. The team considered a stricter fix first, a mandatory second confirmation step on every single recommendation, and rejected it. It would have erased the 29 minutes TraceLine saved on every visit, which defeats the reason the tool exists.
Same fault, same cabinet, replayed with the new card: Priya sees "Corroded connector, 68 to 79 percent" next to "Cracked seal, second most likely." The band is narrow enough to make her check the seal before she opens her toolbox. Four minutes added to her visit. Zero return trips.
I built the single-score version because it tested well in a demo and looked clean on a small screen. It took watching Priya's own wrong-dispatch count nearly quadruple, on a model that hadn't changed at all, to see that the card was never really about looking clean. It was about giving her a reason to keep doing the thing she was already good at.
SPARK, in one screenNot a story about a mistake. SPARK is what forces the design to survive the day it's wrong, before you ship it.
The recap, one line per letter: situation is Priya's 47-minute manual diagnosis, payoff is reading the reasoning instead of the label, anchor is the three-part card, risk is the wrong-dispatch spike the score-only version caused, and keep out is the auto-apply mode the team deliberately didn't ship.
And if you want to be sure it really works, try it somewhere elseSame five letters, a disaster-relief nonprofit instead of a telecom crew. A different flip family shows up here too: this one is about a delegation risk, not over-trust.
Bridge Relief Network runs IntakeLine, an AI tool that reads a family's shelter-intake answers and flags urgent needs for the coordinator on shift. Esteban Duarte manages intake at one of their busiest sites, and used to interview every arriving family himself before IntakeLine existed.
Mapped onto SPARK: situation is Esteban assessing each family by hand on paper forms, asking the same trauma-sensitive questions over and over across shifts. Payoff is coordinators trusting the shared record enough to stop re-asking those questions at every handoff. Anchor is that every auto-suggested flag, like "high priority: medical need," shows the exact answers that triggered it, with the underlying fields editable, not just an accept-or-reject button. Risk is the day IntakeLine misses something because a family answered indirectly; the anchor design lets a coordinator see the raw answers behind the flag and catch the gap immediately, instead of trusting a flag that quietly said nothing was wrong. Keep out is any automatic bed or resource assignment straight from the flag. A person still makes that call.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "show the reasoning, not the label, so trust can be partial," and stop.
Cost: there's no budget this quarter for a full redesign. Say so honestly, and ship the rejected alternative first; it's the cheapest single addition and it catches most of the close calls on its own.
The model gets better, for real: if TraceLine's overall accuracy improves next quarter, that's not a reason to remove the card. A better model still has a worst 5 percent of cases, and those are exactly where a technician needs the reasoning most.
Where people run it wrong.
They treat "transparency" as a disclaimer added after the design is done, not a decision made before it.
They assume more information always helps, and bury the one useful signal, the rejected alternative, under ten fields nobody reads.
They design for the demo case, where the model is right, instead of the case that actually matters: the one where it's wrong.
How to use it live. When someone asks you to design a transparency layer, ask yourself one question first: what does this person do the day the model's top guess is wrong? Design backward from that answer, not forward from what looks clean on a screen.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if the model is right almost every time, doesn't this slow everyone down for nothing?" Response: the four extra minutes only really get spent when the confidence band is already narrow, which is precisely the case where being slow once beats a second truck roll.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Trust, transparency and explainability in UX
- #1 What does a user need to see to trust an AI recommendation?
- #2 Explain the difference between explainability and transparency in a product context.
- #3 How do citations change user behaviour, and what happens when they are wrong?
- #4 Design the disclosure that tells a user they are talking to an AI.
- #5 When does showing the model's reasoning help, and when does it reduce trust?
- #6 Critique a design that surfaces a chain of thought to end users.