InterviewAdvancedShipping & Model Lifecycle / Prototyping with LLMs and rapid POCs / #22

Describe a prototype you would build this week for a problem I name.

The direct answer
Spend the week testing the draft tool against your ugliest real complaints, not your calmest ones, and get two actual reps reading twenty of those drafts by Friday. The prototype has one job this week: find out whether the model's first line, the one naming what this exact person lost, lands as heard or reads as a form letter. Everything else waits until that question has a real answer.
Do this, in order
  1. Test the draft tool against the real worst complaints in the queue this week, not the calm ones, with two reps reading twenty drafts by Friday.Why: a prototype proven only on easy days hasn't proven anything the hard days need.
  2. Make every draft open with one line naming the specific harm, in the customer's own words, and require a rep to confirm that line before anything sends.Why: that one line is where a legitimate angry complaint gets read as heard or gets read as a form letter.
  3. Have two real reps score twenty real drafts, send it, fix it, or bin it, against the actual complaint each one answers.Why: a spec meeting can't tell you whether the tone lands. Only a rep reading it next to the real complaint can.
  4. Watch what a rep does right after they catch one bad draft, not just whether the model was wrong.Why: one tone miss can burn trust in every draft that week, including the good ones, and that costs more time than never building the tool.
  5. Leave multi-language complaints and deciding which ones need a supervisor out of this week entirely.Why: neither question can be answered by a one-week, English-only prototype, and chasing both turns one build into three.
  6. Don't wire the ticketing-system integration or the outage-lookup automation until the tone has passed the worst forty complaints.Why: integration work is real engineering time spent before anyone's confirmed the draft is worth sending at all.

How to answer this, stage by stage

Eight moves. Anchor it to the one line the whole prototype has to survive, not a general pitch for moving fast.

1
Scope it to one concrete prototype
Say it like this
"So I'm the PM for FirstWord at Redbank Power and Light, an AI tool that drafts the first reply to an angry outage complaint. I'm not going to talk about complaint-drafting tools in general. I'll walk through the actual one-week prototype I'd build, tested against real complaints from our last derecho, not a demo built on the calm ones."
Why this works
A scoped example gives the interviewer something to picture and push back on.
2
Say the structure out loud
Say it like this
"Here's how I'll go: why you can't validate this by describing it, the day by day plan for the week, the one decision the whole prototype hangs on, what breaks if I build it too thin or too heavy, and what I'm leaving out of week one on purpose."
Why this works
Two seconds of structure tells the interviewer you have a plan, so they follow instead of guessing where you're going.
3
Reframe what a prototype has to prove
Say it like this
"Most people think you validate a drafting tool by describing what it should do. A mockup, a spec, sign-off in a meeting. That's not really a test. You cannot spec your way to knowing whether a machine sounds heartless to someone whose power has been out for two days. The only way to find out is to put it in front of real, angry language and watch what it actually writes back."
Why this works
This is the actual insight being tested. Skip it and you're just saying "build a prototype," which nobody disagrees with and nobody learns from.
4
Give the anchor
Say it like this
"So here's the one decision the whole prototype hangs on. Every draft opens with a single line naming the specific thing this person said they lost, close to their own words, not a category. Not 'your outage.' 'Your dad's oxygen machine.' A rep has to read and confirm that exact line before the rest of the draft even gets read, and nothing sends without them."
Why this works
Naming the specific mechanism, out loud, is what separates a real design decision from a vague "make it empathetic" gesture.
5
Walk the actual week, day by day
Say it like this
"Day one, I pull two hundred real complaints from our last three outage events, sorted by how angry they actually are, in their own words, not ones I write myself. Days two and three, I wire the generator so it reads the complaint plus the real outage data for that address and writes the harm-first draft. Day four, I run it against the forty worst complaints in that pile, the ones with words like 'medical' and 'spoiled' and 'my mother,' not the polite ones asking when the power's back. Day five, two real reps read twenty of those drafts blind, next to what they'd have written themselves, and mark each one: send it, fix it, or bin it."
Why this works
A day-by-day plan proves you've actually thought through the constraint of one week, not just gestured at "build fast."
6
Prove it with what the test found
Say it like this
"Here's what day five found. Nine of the forty drafts flattened the harm. One called a customer's oxygen concentrator 'your equipment.' A rep almost sent it before catching it in the blind read. But here's the part that mattered more. After reading two or three bad drafts in a row, both reps stopped trusting any of the tool's drafts, even the easy thirty, and started rewriting every single one from scratch anyway, just in case. Reply time went from twelve minutes by hand to fifteen minutes with the tool sitting right there."
Why this works
A real number, with the real cost attached, does more work than "the model wasn't perfect" ever will.
7
Say what's kept out of week one
Say it like this
"For week one, I'm not touching two things. No other languages, and no deciding which complaints need to go to a supervisor. Both of those need their own week. Trying to prove three things at once this week means proving none of them."
Why this works
Naming what's out shows judgment instead of a wish list, and protects the one question this week actually has time to answer.
8
Close on the number
Say it like this
"So: test the harm-first draft against the ugliest real complaints, not the calm ones, two reps reading twenty of them by Friday. Nine of forty drafts missed the harm this week, and that number matters less than what happened right after it. Reply time went from twelve minutes to fifteen, worse than never building the thing, until the reps could trust the line again."
Why this works
Interviewers remember the last line most, and this one hands them something they can check, not just a mood.

Let's learn

What happens the first time an AI tries to write back to someone whose father's oxygen machine just lost power?

Redbank Power and Light built FirstWord: read an angry outage complaint, check the real outage data for that address, and write the first draft of the reply, so a rep edits it instead of writing it from scratch.

Hand-sketch with a person icon at the center labeled Anwar, complaint desk, and four labeled callouts around it: printed complaint emails, reply written by hand, paper outage map, red pen for the worst ones
Today, without FirstWord

Before FirstWord, a rep read the complaint, pulled the outage record for that address by hand, and typed a reply. About 12 minutes each. On a normal week that's fine. On the second Tuesday of June, a derecho knocked out power to 14,000 homes in Redbank's service area, and the complaint queue hit 900 emails in the first day and a half. At 12 minutes a reply, some customers waited over 30 hours to hear back at all.

Knowledge spark: what's a prototype actually for? A prototype isn't a smaller version of the finished product. It's a fast way to find out if one risky guess is true, before you spend real weeks building around it.

Anwar Bellweather, FirstWord's PM, didn't ask for a spec review before the build. He asked for one week: pull real complaints, wire a draft generator, and put it in front of two reps by Friday.

On the routine complaints, the thirty polite ones just asking when the power's coming back, it looked great. A draft in about 40 seconds, a rep needed about 90 more to check and send it. Two and a half minutes, against 12 by hand.

Then day four ran it against the 40 worst complaints in the pile, the ones with "medical" and "spoiled" and "my mother" in them. Nine of forty drafts flattened the harm. One called a customer's father's oxygen concentrator "your equipment."

Minutes to send a first reply, three points in the week
16 min 0 min 8 min 12 min before FirstWord written by hand 2.5 min easy complaints with FirstWord 15 min after trust broke every reply, rewritten anyway
the old, manual number the draft working the draft after trust broke
The tool cut reply time by more than three quarters on the easy complaints. Then nine bad drafts on the hard ones pushed the whole desk past where it started.
Nine bad drafts didn't cost the reps nine replies. It cost them all forty.

At its worst, this doesn't stay small. Redbank's complaint desk had 900 emails backed up that week. If every one of them went from 12 minutes to 15 because reps stopped trusting any draft, that's 2,700 extra minutes, about 45 hours, spread across a six-person desk in the exact week they could least afford it. And the actual problem, one line that named the wrong thing, would still be sitting there unfixed.

The decision that mattered Test the prototype against the real worst complaints first, not the calm ones saved for later in the week. A tool that only survives easy days has proven nothing about the day it actually needs to work.

The choice I would take back. On day one, the plan was to wire the draft generator against the calmest complaints first, because they were fastest to build cleanly and easiest to show leadership by midweek. The worst ones got pushed to day four, almost as an afterthought. That made sense when the goal was proving the pipeline ran end to end. It stopped making sense the moment the real question became whether the tone survives real anger, not whether the code runs.

What I would leave alone. The routine complaints, the ones just asking when the power's back, with no medical word or spoiled freezer or lost day of pay in them, never needed the harm-first line to begin with. A plain "thanks for reaching out, here's your update" already worked fine for those, in every test all week.

The lesson. A prototype that only proves itself on the easy days hasn't proven anything yet. The whole reason to build it is for the hard ones.

Now here is the same thing as a story

The short version is above. Read this one when you've got a few minutes, for why it mattered.

The complaint desk at Redbank gets loud around hour six of a big outage, when the phone calls turn into emails and the emails turn into all capital letters.

Anwar Bellweather had run product for the customer-care team for three years, long enough to know that most outage complaints are the same five sentences with different names in them. He also knew the ones that weren't. On the second Tuesday of June, a derecho tore through Redbank's service area and cut power to 14,000 homes. By the next afternoon, 900 complaints sat in the queue, and the six reps on the desk were 30 hours behind on the first one.

Leadership wanted a plan for an AI drafting tool. Anwar asked for something smaller first: one week, real complaints, two reps, no roadmap yet.

Day one, he pulled 200 real complaints from the last three outage events, not ones anyone wrote to look good in a demo. Sorted by how angry they actually read. Days two and three, an engineer wired the draft generator: feed it the complaint and the real outage data for that address, and it would write back.

Hand-sketch flow of four boxes connected left to right: angry email, find what's lost, name it first highlighted in green as the anchor step, rep confirms and sends
The anchor: name the harm before anything else

Anwar made one rule non-negotiable before a single draft got tested. The first line of every draft had to name the specific thing that person said they lost, in close to their own words. Not "your outage." "Your dad's oxygen machine." And a rep had to read and confirm that exact line before the rest of the draft counted for anything. Nothing would send on its own.

Day four, they ran it against the 40 worst complaints in the pile. The ones that mentioned a medical device, spoiled food, a missed shift, an elderly parent home alone. Not the calm thirty asking when the power was back.

Nine of forty flattened the harm. One read the words "my dad's oxygen machine went off Tuesday night" and drafted back: "We're sorry your equipment was affected by the outage. Our crews are working as fast as they can."

On day five, during the blind read, a rep had that exact draft open, cursor over the send button, three complaints deep into a stack of twenty. She almost sent it. Something about the word "equipment" made her stop and reread the original email first. That was the whole difference. One second of pause, on one draft, out of forty.

Nine bad drafts didn't cost the reps nine replies. It cost them all forty.

Because here's what happened after. Both reps had now seen two or three drafts in a row that missed the point in a way that felt, to them, careless with someone's actual fear. So they stopped trusting any of FirstWord's drafts. Not just the nine bad ones. All forty, including the thirty-one that were fine. They went back to reading each complaint in full and writing the reply from scratch, the way they always had, except now they were also reading a draft first out of habit before throwing it away. Average reply time for the batch: 15 minutes. Three minutes worse than before FirstWord existed at all.

Two-panel hand-sketch comparison. Left panel, a question mark icon labeled model misreads it, caption calls dad's oxygen machine your equipment. Right panel, a document icon labeled rep catches it, caption rewrites the one line before it sends
The day it's wrong, and the anchor still catches it

So here is the decision Anwar took back.

On day one, the plan was to build against the calm complaints first, because they were faster to wire cleanly and easier to show around by Wednesday. The worst forty got pushed to day four, almost as an afterthought, something to squeeze in if there was time. That made sense when the goal was proving the pipeline worked at all. It stopped making sense the second the real question became whether the tone survives someone's actual worst day, not whether the code compiles.

Anwar didn't scrap FirstWord. He rewrote the test order. The next round put the worst 40 first, on day two instead of day four, before either rep had spent a single hour trusting the easy ones. Of the next 40 hard complaints tested, 3 flattened the harm instead of 9, and both misses got caught in the confirm step before anything sent.

And the thing I'd want to tell myself, back when the plan was to save the hard complaints for later in the week: a tool only earns trust on the day it survives the worst thing someone says to it. Every day before that is just rehearsal.

SPARK, run against an angry inbox

This question sounds like it wants a metric, how would you measure success, or a process answer, how would you roll this out. It's really asking for one concrete design decision that has to survive contact with real anger inside a fixed week, so SPARK fits. A question asking how Anwar would know FirstWord was working three months into rollout would reach for LEAD instead.

S, situation. Today, without FirstWord, a rep reads an angry outage email, checks the outage system by hand for the cause and the crew's ETA, and writes the reply from nothing. No model has attempted this on real angry language yet.
P, payoff. Not "save reps time." The habit worth building this week: find out whether the model's tone actually lands on real anger, before a quarter gets spent on specs and integrations nobody's confirmed will work.
A, anchor. Every draft opens with one line naming the specific harm, in the customer's own words, and a rep must confirm that exact line before anything reads further or sends.
R, risk. Too thin, and the prototype only gets tested on calm complaints, ships looking great, then breaks on the first real bad one. Too heavy, and the week gets spent wiring the ticketing integration and the outage lookup before anyone's confirmed the tone works at all.
K, keep out. Multi-language complaints, and deciding which complaints need to route to a supervisor. Both are real problems and neither can be answered by a one-week, English-only prototype.
Two-panel hand-sketch comparison labeled what we left for later. Left panel, a document icon labeled day one, caption harm-first draft, English, a rep reads every line. Right panel, a greyed question mark icon labeled not day one, caption other languages, deciding who needs a supervisor
What we left for later, kept visibly separate from day one
Why the anchor survives the risk Check it against the near miss. Does the harm-first line still teach you something even when the model gets the harm wrong? Yes, because the confirm step is built to catch exactly that miss before it sends, not after. Does it avoid the over-build trap? Yes, because K keeps two entire other problems explicitly off this week's job.

And if you want to be sure it really works, try it somewhere else

A diagnostic imaging clinic runs on a completely different clock, but the same gap between what a machine writes and what a scared person reads shows up in a radiology report.

S. Odell Radke runs product at Hawthorn Imaging. Today, without a prototype, a patient opens their radiology report in the portal, hits a phrase like "hypoattenuating lesion," and calls the clinic asking if it's serious, usually afraid.
P. The habit worth building: find out this week whether one added plain-language line actually stops the scared callback, before Hawthorn commits to rewriting how every report reaches a patient.
A. Same shape, a different desk. Odell tests a prototype that adds one line above the real report: what was found, and whether it needs follow-up soon or not, drafted from the report's own findings, checked by a radiologist before a patient ever sees it.
R. Test it on too few reports over one quiet week, and a good result looks like proof it's ready for findings it's never seen. Try to also build the full glossary and the follow-up scheduling link in the same week, and the one real question, does the line stop the callback, never gets a clean answer.
K. No full glossary of medical terms yet, and no automatic booking of follow-up visits. Both are later decisions, once the one-line summary is proven to work at all.

Minutes before a patient understands their result, before and after one line
25 min 0 min before the line 25 min, on average after the line 4 min, on average
reading the report looking up the jargon on hold, calling the clinic
Reading time barely changes. The jargon lookup and the phone call, the two things one honest sentence can answer directly, almost disappear.

On the 30 real reports with the most historically confusing wording, patients used to spend about 25 minutes total: 2 reading, 15 looking up terms, 8 on hold waiting to ask the clinic if it was serious. With the one-line summary tested in front of them first, that dropped to about 4 minutes, because the line answered the one question the call was always really about.

Swap the trigger and it still runs

  • Speed: even if FirstWord wrote every draft instantly instead of in a few seconds, that wouldn't fix a draft that names the wrong thing as lost. Speed and accuracy are two different jobs.
  • Cost: if wiring the draft generator cost nothing at all, that still wouldn't tell you which nine of forty complaints the tone actually misses. You still need real angry language, not free engineering time.
  • The model gets better: if FirstWord's grammar and phrasing became flawless, that still wouldn't fix whether it named the right loss, because polish and accuracy are two different jobs too.

Where people run it wrong

  • Testing the draft generator only on the calm, easy complaints because they're faster to wire, and calling that proof it's ready for the ugly ones.
  • Letting drafts auto-send once the pipeline runs cleanly, before any rep has read one against the real complaint it answers.
  • Trying to also settle escalation routing or multi-language support inside the same one-week build, so the real question, does the tone survive anger, never gets a clean answer.

How to use it live

If you're asked this cold, ask yourself what the single worst real example in your own inbox or ticket queue looks like, then design the prototype's first test around that one, not a clean example you made up. That question, asked of yourself out loud, finds the real test faster than trying to write the perfect week-one plan.

Flashcards (click a card to flip it)

1 · THE SITUATION
What's the situation, before this prototype existed?
Tap to flip
ANSWER
Redbank reps wrote every outage-complaint reply from scratch, checking the outage system by hand, at about 12 minutes each, with no AI having tried it on real angry language.
2 · THE PAYOFF
What's the real habit this one-week prototype is trying to build?
Tap to flip
ANSWER
Finding out this week whether the model's tone actually lands on real angry complaints, before a quarter gets spent on specs and integrations nobody's confirmed will work.
3 · THE ANCHOR
What's the one design decision the whole prototype hangs on?
Tap to flip
ANSWER
Every draft opens with one line naming the specific thing the customer said they lost, in their own words, and a rep must confirm that line before anything sends.
4 · THE RISK
What breaks if the prototype runs too thin, or too heavy for one week?
Tap to flip
ANSWER
Too thin, and it's tested only on calm complaints, ships, and breaks on the first real bad one. Too heavy, and the week gets spent wiring integrations before anyone's confirmed the tone works at all.
5 · THE PROOF
What did the day-four test find that a spec review never would have?
Tap to flip
ANSWER
Nine of forty drafts flattened the harm, one calling a customer's father's oxygen concentrator "your equipment." A rep nearly sent it before catching it in the blind read.
6 · THE NUMBER
___ of ___ drafts flattened the harm on the worst complaints, and reply time went from ___ minutes to ___ once trust broke.
Tap to flip
ANSWER
9 of 40 flattened the harm. Reply time went from 12 minutes to 15.
7 · THE REPLAY
Same bad storm, worst complaints tested first instead of last. What changes?
Tap to flip
ANSWER
Of the next 40 hard complaints, 3 flattened the harm instead of 9, and both misses got caught in the confirm step. Reps kept trusting the other drafts instead of rewriting everything from scratch.
8 · CROSS-PRODUCT
Section 4 runs SPARK again on a different product. Which one, and what does its anchor test?
Tap to flip
ANSWER
Hawthorn Imaging's plain-language radiology cover line. Its anchor tests whether a patient can tell if a result needs follow-up soon, from the first line, without calling the clinic.

Check yourself Score: 0 / 0

True or false
1. True or false: the nine drafts that flattened the harm were the real cost of testing the prototype on the worst complaints.
  • True
  • False
Show hint
Look at what the reps actually did after reading two or three bad drafts in a row.
Show answer
False. The real cost was the reps losing trust in all forty drafts, including the thirty-one good ones, and going back to writing every reply from scratch, which pushed reply time past the original twelve minutes.
Fill in the blank
2. ___ of ___ drafts flattened the harm on the worst complaints, and reply time went from ___ minutes to ___ once the reps stopped trusting the tool.
Show hint
This number shows up twice, once in the story, once in the chart.
Show answer
9 of 40; 12 minutes to 15. Fewer than a quarter of the hard drafts missed, but the cost of those misses landed on every reply that week, not just the nine.
Multiple choice
3. Which design matches the anchor this answer argues for?
  • A. A generic empathy template added to the top of every reply.
  • B. A first line naming the specific harm, in the customer's own words, that a rep must confirm before anything sends.
  • C. Auto-sending any draft that scores below a set anger threshold.
  • D. A longer spec document, reviewed by legal, before any code gets written.
Show hint
The anchor needs the exact harm named, plus a human check before anything reaches the customer.
Show answer
B. A is generic, not specific to what this person actually lost. C removes the human check that catches the near miss. D is still description, not a real test.
Short answer
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Think about why testing on the calm complaints first sounded fine, back before the real question was known.
Show answer
Model answer: The plan was to wire the draft generator against the calmest complaints first, because they were fastest to build cleanly and easiest to show around by midweek. That made sense when the goal was proving the pipeline ran end to end. It stopped making sense once the real question became whether the tone survives real anger, not whether the code runs.
Short answer, apply it yourself
5. Pick a product you use yourself, or one your team is building. What's the one risky guess a one-week prototype could test on real, ugly examples, instead of a mockup or a spec?
Show hint
Look for the guess that would embarrass you most if it turned out wrong after launch, not the easiest one to test.
Show answer
Model answer: "Our support bot drafts replies to refund requests. The risky guess is whether it can tell a customer who's genuinely out of pocket from one who's fishing for a freebie. A week testing it on our twenty angriest real refund threads, with a human reading every draft, would show that faster than any policy document would."
Multiple choice
6. Based on this answer's own numbers, if only 2 of 40 drafts had flattened the harm instead of 9, would testing the worst complaints first still have been the right call?
  • A. Yes, because the point of testing the hard complaints first is finding out whether the tone survives them at all, and that's true whether the miss rate is high or low.
  • B. No, at 2 of 40 the team should have skipped the confirm step and let drafts send automatically.
  • C. No, a lower miss count means the harm-first line wasn't necessary in the first place.
  • D. Yes, but only because a lower number would have looked better in front of leadership.
Show hint
Compare what a lower miss count changes about the value of testing the hard cases first, against what it changes about whether a human still needs to check.
Show answer
A. A lower miss rate doesn't remove the need to know it, or the need for a rep to catch the ones that remain. It would still have been the wrong week to find that out on the calm complaints instead.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more