Describe a prototype you would build this week for a problem I name.
- Test the draft tool against the real worst complaints in the queue this week, not the calm ones, with two reps reading twenty drafts by Friday.Why: a prototype proven only on easy days hasn't proven anything the hard days need.
- Make every draft open with one line naming the specific harm, in the customer's own words, and require a rep to confirm that line before anything sends.Why: that one line is where a legitimate angry complaint gets read as heard or gets read as a form letter.
- Have two real reps score twenty real drafts, send it, fix it, or bin it, against the actual complaint each one answers.Why: a spec meeting can't tell you whether the tone lands. Only a rep reading it next to the real complaint can.
- Watch what a rep does right after they catch one bad draft, not just whether the model was wrong.Why: one tone miss can burn trust in every draft that week, including the good ones, and that costs more time than never building the tool.
- Leave multi-language complaints and deciding which ones need a supervisor out of this week entirely.Why: neither question can be answered by a one-week, English-only prototype, and chasing both turns one build into three.
- Don't wire the ticketing-system integration or the outage-lookup automation until the tone has passed the worst forty complaints.Why: integration work is real engineering time spent before anyone's confirmed the draft is worth sending at all.
How to answer this, stage by stage
Eight moves. Anchor it to the one line the whole prototype has to survive, not a general pitch for moving fast.
Let's learn
What happens the first time an AI tries to write back to someone whose father's oxygen machine just lost power?
Redbank Power and Light built FirstWord: read an angry outage complaint, check the real outage data for that address, and write the first draft of the reply, so a rep edits it instead of writing it from scratch.
Before FirstWord, a rep read the complaint, pulled the outage record for that address by hand, and typed a reply. About 12 minutes each. On a normal week that's fine. On the second Tuesday of June, a derecho knocked out power to 14,000 homes in Redbank's service area, and the complaint queue hit 900 emails in the first day and a half. At 12 minutes a reply, some customers waited over 30 hours to hear back at all.
Anwar Bellweather, FirstWord's PM, didn't ask for a spec review before the build. He asked for one week: pull real complaints, wire a draft generator, and put it in front of two reps by Friday.
On the routine complaints, the thirty polite ones just asking when the power's coming back, it looked great. A draft in about 40 seconds, a rep needed about 90 more to check and send it. Two and a half minutes, against 12 by hand.
Then day four ran it against the 40 worst complaints in the pile, the ones with "medical" and "spoiled" and "my mother" in them. Nine of forty drafts flattened the harm. One called a customer's father's oxygen concentrator "your equipment."
At its worst, this doesn't stay small. Redbank's complaint desk had 900 emails backed up that week. If every one of them went from 12 minutes to 15 because reps stopped trusting any draft, that's 2,700 extra minutes, about 45 hours, spread across a six-person desk in the exact week they could least afford it. And the actual problem, one line that named the wrong thing, would still be sitting there unfixed.
The choice I would take back. On day one, the plan was to wire the draft generator against the calmest complaints first, because they were fastest to build cleanly and easiest to show leadership by midweek. The worst ones got pushed to day four, almost as an afterthought. That made sense when the goal was proving the pipeline ran end to end. It stopped making sense the moment the real question became whether the tone survives real anger, not whether the code runs.
What I would leave alone. The routine complaints, the ones just asking when the power's back, with no medical word or spoiled freezer or lost day of pay in them, never needed the harm-first line to begin with. A plain "thanks for reaching out, here's your update" already worked fine for those, in every test all week.
The lesson. A prototype that only proves itself on the easy days hasn't proven anything yet. The whole reason to build it is for the hard ones.
Now here is the same thing as a story
The short version is above. Read this one when you've got a few minutes, for why it mattered.
The complaint desk at Redbank gets loud around hour six of a big outage, when the phone calls turn into emails and the emails turn into all capital letters.
Anwar Bellweather had run product for the customer-care team for three years, long enough to know that most outage complaints are the same five sentences with different names in them. He also knew the ones that weren't. On the second Tuesday of June, a derecho tore through Redbank's service area and cut power to 14,000 homes. By the next afternoon, 900 complaints sat in the queue, and the six reps on the desk were 30 hours behind on the first one.
Leadership wanted a plan for an AI drafting tool. Anwar asked for something smaller first: one week, real complaints, two reps, no roadmap yet.
Day one, he pulled 200 real complaints from the last three outage events, not ones anyone wrote to look good in a demo. Sorted by how angry they actually read. Days two and three, an engineer wired the draft generator: feed it the complaint and the real outage data for that address, and it would write back.
Anwar made one rule non-negotiable before a single draft got tested. The first line of every draft had to name the specific thing that person said they lost, in close to their own words. Not "your outage." "Your dad's oxygen machine." And a rep had to read and confirm that exact line before the rest of the draft counted for anything. Nothing would send on its own.
Day four, they ran it against the 40 worst complaints in the pile. The ones that mentioned a medical device, spoiled food, a missed shift, an elderly parent home alone. Not the calm thirty asking when the power was back.
Nine of forty flattened the harm. One read the words "my dad's oxygen machine went off Tuesday night" and drafted back: "We're sorry your equipment was affected by the outage. Our crews are working as fast as they can."
On day five, during the blind read, a rep had that exact draft open, cursor over the send button, three complaints deep into a stack of twenty. She almost sent it. Something about the word "equipment" made her stop and reread the original email first. That was the whole difference. One second of pause, on one draft, out of forty.
Because here's what happened after. Both reps had now seen two or three drafts in a row that missed the point in a way that felt, to them, careless with someone's actual fear. So they stopped trusting any of FirstWord's drafts. Not just the nine bad ones. All forty, including the thirty-one that were fine. They went back to reading each complaint in full and writing the reply from scratch, the way they always had, except now they were also reading a draft first out of habit before throwing it away. Average reply time for the batch: 15 minutes. Three minutes worse than before FirstWord existed at all.
So here is the decision Anwar took back.
On day one, the plan was to build against the calm complaints first, because they were faster to wire cleanly and easier to show around by Wednesday. The worst forty got pushed to day four, almost as an afterthought, something to squeeze in if there was time. That made sense when the goal was proving the pipeline worked at all. It stopped making sense the second the real question became whether the tone survives someone's actual worst day, not whether the code compiles.
Anwar didn't scrap FirstWord. He rewrote the test order. The next round put the worst 40 first, on day two instead of day four, before either rep had spent a single hour trusting the easy ones. Of the next 40 hard complaints tested, 3 flattened the harm instead of 9, and both misses got caught in the confirm step before anything sent.
And the thing I'd want to tell myself, back when the plan was to save the hard complaints for later in the week: a tool only earns trust on the day it survives the worst thing someone says to it. Every day before that is just rehearsal.
SPARK, run against an angry inbox
This question sounds like it wants a metric, how would you measure success, or a process answer, how would you roll this out. It's really asking for one concrete design decision that has to survive contact with real anger inside a fixed week, so SPARK fits. A question asking how Anwar would know FirstWord was working three months into rollout would reach for LEAD instead.
And if you want to be sure it really works, try it somewhere else
A diagnostic imaging clinic runs on a completely different clock, but the same gap between what a machine writes and what a scared person reads shows up in a radiology report.
S. Odell Radke runs product at Hawthorn Imaging. Today, without a prototype, a patient opens their radiology report in the portal, hits a phrase like "hypoattenuating lesion," and calls the clinic asking if it's serious, usually afraid.
P. The habit worth building: find out this week whether one added plain-language line actually stops the scared callback, before Hawthorn commits to rewriting how every report reaches a patient.
A. Same shape, a different desk. Odell tests a prototype that adds one line above the real report: what was found, and whether it needs follow-up soon or not, drafted from the report's own findings, checked by a radiologist before a patient ever sees it.
R. Test it on too few reports over one quiet week, and a good result looks like proof it's ready for findings it's never seen. Try to also build the full glossary and the follow-up scheduling link in the same week, and the one real question, does the line stop the callback, never gets a clean answer.
K. No full glossary of medical terms yet, and no automatic booking of follow-up visits. Both are later decisions, once the one-line summary is proven to work at all.
On the 30 real reports with the most historically confusing wording, patients used to spend about 25 minutes total: 2 reading, 15 looking up terms, 8 on hold waiting to ask the clinic if it was serious. With the one-line summary tested in front of them first, that dropped to about 4 minutes, because the line answered the one question the call was always really about.
Swap the trigger and it still runs
- Speed: even if FirstWord wrote every draft instantly instead of in a few seconds, that wouldn't fix a draft that names the wrong thing as lost. Speed and accuracy are two different jobs.
- Cost: if wiring the draft generator cost nothing at all, that still wouldn't tell you which nine of forty complaints the tone actually misses. You still need real angry language, not free engineering time.
- The model gets better: if FirstWord's grammar and phrasing became flawless, that still wouldn't fix whether it named the right loss, because polish and accuracy are two different jobs too.
Where people run it wrong
- Testing the draft generator only on the calm, easy complaints because they're faster to wire, and calling that proof it's ready for the ugly ones.
- Letting drafts auto-send once the pipeline runs cleanly, before any rep has read one against the real complaint it answers.
- Trying to also settle escalation routing or multi-language support inside the same one-week build, so the real question, does the tone survive anger, never gets a clean answer.
How to use it live
If you're asked this cold, ask yourself what the single worst real example in your own inbox or ticket queue looks like, then design the prototype's first test around that one, not a clean example you made up. That question, asked of yourself out loud, finds the real test faster than trying to write the perfect week-one plan.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Prototyping with LLMs and rapid POCs
- #1 What can you learn from a prototype that you cannot learn from a spec?
- #2 Describe how you would build a working prototype of an AI feature in a day.
- #3 What are the risks of a PM prototyping without engineering involvement?
- #4 Explain when a Wizard of Oz prototype beats a real model.
- #5 How do you keep a prototype from setting unrealistic expectations?
- #6 Describe the difference between a demo prototype and a learning prototype.