CaseIntermediateModel Fluency & the AI PM Role / Working with ML engineers and researchers / #2

An engineer says the model cannot do that. What questions do you ask before accepting it?

TRACE · what an engineer's no actually means, tested on a shelved opener feature at Kerning

Kerning drafts opening lines and reviews message drafts for people navigating dating apps, and Auden Colter owns the roadmap for how personal those openers actually feel. This is the four months a promising idea sat shelved because one honest, one afternoon test got written into the plan as the final word.

The direct answer
Don't accept "the model can't do that" off one report. Ask exactly what got tried: a single plain instruction, or real examples of the target voice sitting in the prompt. Ask what "reliably enough" would actually mean and whether that bar got tested, and whether the task ever got broken into smaller steps. Then ask the one question that decides everything: has anyone run a real test against this exact case, or is this a belief about how models behave in general.
Do this, in order
  1. Ask exactly what got tried before writing "can't" into the plan.Why: a single zero-shot instruction and a real test with target-voice examples are different pieces of evidence, and only one of them settles anything.
  2. Ask what "reliably enough" would mean, and whether that bar got tested.Why: "the model can't do X" often really means "it missed the bar once," a smaller claim wearing a bigger one's clothes.
  3. Ask whether a workaround, like splitting the task into steps, got tried.Why: some tasks fail in one shot and land the moment the hard part gets pulled out and generated against on its own.
  4. Ask for the actual prompt and output as evidence, not a verbal summary.Why: a claim with no test attached is a belief wearing a fact's clothes, and it reads exactly like a real wall eighteen months later.
  5. Log every "can't" claim with exactly what was tried.Why: without a record, nobody can tell a settled wall from an untested guess when the question comes back around.
  6. Don't reopen a claim on a hunch, only on a genuinely new approach worth testing.Why: real capability walls exist too, and re-checking a well-tested no wastes the same time this whole method is meant to save.

How to answer this, stage by stage

Nobody is grading whether you know what a token is. They're grading whether you'll let one afternoon's test decide a roadmap, or ask for the receipt first.

1
Scope it to one team, one claim
Say it like this
"Let's ground this in one real case. Kerning drafts opening lines and reviews message drafts for people on dating apps. A product lead pitched openers that actually sound like the user, in their own voice. A senior engineer said the model can't do that. That's the claim I'm going to pressure test."
Why this works
One real feature and one real claim stops the answer turning into a lecture on prompting in the abstract.
2
Say your structure out loud
Say it like this
"I'll run this as TRACE. Timeline: what got tried, and when. Recut: what 'cannot' could actually mean, there's more than one option. Assume nothing: check which meaning we're actually dealing with. Cause candidates: the real questions I'd ask. Evidence test: the one question that separates a wall from a guess."
Why this works
Two seconds of structure tells the interviewer you have a method, not a hunch dressed up as confidence.
3
Reframe the question: "cannot" is not one claim
Say it like this
"Before I ask anything, I want to flag something. 'The model can't do that' gets used for at least four different situations: a genuine hard wall, an approach nobody's tried yet, something the model can do but not reliably, or 'I don't want to spend the time finding out.' Treating all four the same is the actual mistake here."
Why this works
This is the recut. It's the move that separates a real diagnosis from a candidate who just lists questions to ask.
4
Give the decision, committed
Say it like this
"So here's what I'd actually do. I would not write this into the plan off one report. I'd ask what exactly got tried, what 'reliably' would need to mean and whether that bar got tested, and whether anyone tried breaking the task into steps."
Why this works
This is the direct answer to the question, said plainly, before a single detail of the story shows up.
5
Prove it with the failure, numbers first
Say it like this
"Here's what actually happened at Kerning. An engineer spent one afternoon on three zero-shot prompts, got a generic opener every time, and wrote 'the model can't do this' into the plan. It got shelved. Four months later, a different engineer, working on something else, fed the model five real examples of the target voice in the same prompt. It worked in about twenty minutes."
Why this works
A real, checkable failure beats a hypothetical every single time an interviewer hears one.
6
Name the evidence test
Say it like this
"The question that would have caught this in week one: 'Did you test it with real examples of the voice sitting in the prompt, or did you just describe the voice and ask for it?' Nobody asked that. If they had, the honest answer was 'I only tried the second one.'"
Why this works
This is the strongest move in TRACE, and it's checkable against something real, not a guess about which explanation sounds better.
7
Say the trade-off and what you'd leave alone, then close
Say it like this
"One thing worth saying plainly: the fix costs more, roughly three times the tokens and a bit more time per opener, since five examples ride along with every call. I'd take that trade. And one place I wouldn't go pressure-testing every 'can't': a claim the team has already tested three different ways with real evidence attached. So, to close it: ask what was tried before you accept what wasn't possible."
Why this works
Naming the cost and the boundary shows judgment on both sides, and the close restates the decision in one breath.

Let's learn

Kerning reads a person's dating profile and a match's profile, then drafts an opening message and checks message drafts before they get sent, so starting a chat stops feeling like homework.

For its first two years, every opener Kerning wrote followed one recipe: find a topic in the match's profile, mention it, ask a light question. It worked well enough. Reply rates on those openers held around 22 percent, in a two week test cohort, close to what people got writing their own generic openers by hand. Users kept using it. But the single most repeated line in Kerning's own feedback box was some version of "it works, but it doesn't sound like me."

Knowledge spark: zero-shot versus few-shot Zero-shot means one plain instruction, no examples, just "write it like this." Few-shot means the same instruction, plus a handful of real examples of exactly what good output looks like, sitting right there in the same prompt.

A product lead pitched a fix: openers written in the user's own voice, not just about the user's own topic. A senior engineer tried it, one plain instruction, no examples, three different times. Each time the model handed back the same kind of line it always wrote: pleasant, topical, generic. So the engineer reported back: the model can't do this. It just defaults to a generic compliment no matter how you ask.

Here is the turn. That report was true, and it was also not proof of anything close to what it got used for.

The three tries that failed weren't proof the model couldn't do this. They were proof that one way of asking didn't work.

The idea got written into the plan as "not feasible: a model limitation" and moved to the bottom of the backlog. Nobody attached the actual prompts that were tried. Nobody wrote down what "sounds like someone's voice" would even need to look like to count as working. The sentence "the model can't do voice matching" just started getting repeated in planning meetings, the way a fact gets repeated once nobody remembers it was ever a guess.

Hand sketched comparison diagram titled Filed as fact, not touched for four months. Left panel, a question mark icon labeled Filed: not feasible, caption one afternoon, three zero-shot tries. Right panel, a box icon labeled Sitting untouched, caption four months, complaint theme holds at 18 percent.
One afternoon's test became a permanent line in the plan. Nobody circled back to check it for four straight months.

What it costs at its worst: four straight months where the single most common complaint kept showing up in almost one out of every five pieces of critical feedback, a live, named, understood problem, sitting completely untouched, because everyone believed it had already been ruled out.

The choice I would take back The engineer's honest report from one afternoon, with three tries, got treated as the final word on the whole idea, instead of one data point that still needed a receipt attached: the actual prompt that was tried, and the actual output it produced.

What I would leave alone: Kerning genuinely cannot know whether a match is going to reply. That's not a prompting gap. No number of examples fixes a question about someone else's future choice, so that kind of "can't" doesn't need four sharp questions, it needs to be accepted and designed around instead.

The lesson: a "no" backed by one afternoon and a "no" backed by three genuinely different approaches use the exact same word, but they are not the same kind of no. My job isn't to doubt every engineer's read. It's to make sure "can't" always comes with a receipt before it becomes a fact everyone repeats.

Now here is the same thing as a story

Say the short version out loud in an interview. Read this one when you want to feel exactly how one true sentence, said once, quietly became a fact nobody checked again for four months.

Rieko Cleary can read a bad model output and tell you, inside one sentence, whether the problem is the prompt or the model itself. Most engineers guess. She checks. In three years on Kerning's small AI team, she's the person people go to when a feature isn't behaving and nobody knows why.

Every Tuesday at ten, the product team crowds into the same glass conference room to walk the roadmap, and by the second hour the air always turns thick. In February, Auden Colter, who owns how personal Kerning's openers actually feel, brought a new idea into that room: openers written in a user's own voice, not just about the user's own topic. Someone dry and self deprecating in their bio should get an opener that sounds dry and self deprecating. Someone warm and full of exclamation points should get one that sounds warm.

Rieko liked the idea and said she'd try it that week.

She spent one afternoon on it. Three prompts, each a version of the same plain instruction: "Write an opening line in this person's voice, based on their bio, that references something specific about their match." Three times, the model handed back a pleasant, topical line that could have come from anyone's bio. Nothing dry. Nothing warm. Just generic, the way Kerning's openers had always been generic.

She wrote it up honestly, the way she writes up everything. "Tried it three different ways this afternoon. The model just defaults to a generic compliment, no matter what's in the bio. I don't think it can actually pick up on someone's specific voice from a paragraph of text."

Auden read that Thursday, before the next roadmap review. It sounded right. It sounded like the kind of thing someone who'd actually tried it would say. Voice Match got marked "not feasible: a model limitation" and moved to the bottom of the icebox, and the room moved on to the next item on the list.

Nobody wrote down the three prompts Rieko had actually tried. Nobody wrote down what "sounds like someone's voice" would need to look like to count as working. The sentence "the model can't do voice matching" just started getting repeated in planning meetings, the way a fact gets repeated once nobody remembers it was ever a guess.

For four straight months, nothing about it changed. Kerning's own feedback tool kept tagging the same complaint theme, some version of "it works, but it doesn't sound like me," and it held between 17 and 19 percent of all critical feedback every single month, February through May. Auden saw the number every month. It never occurred to her to connect it back to a feature the team had already ruled out.

Hand sketched horizontal timeline titled Three months of silence, then twenty minutes. Four marks along the line. Tuesday review, caption Rieko: the model can't do this. Filed: not feasible, caption Voice Match shelved. Four months untouched, caption complaint theme holds near 18 percent. Meara's retest, this mark emphasized in coral, caption five voice examples, twenty minutes.
One afternoon in February became four flat months. The gap between the claim and the retest is the entire point.

In June, Meara Osei-Tutu, who'd joined the AI team two months earlier, was building something unrelated: a tone checker meant to flag when a user's draft reply didn't match the tone of the conversation so far. To teach it what "dry," "warm," and "flirty" actually looked like, she fed the model five real labeled examples of each, pulled straight from real message threads, sitting right there in the prompt.

Almost by accident, she flipped one of her own test prompts around. Instead of asking the model to label a message's tone, she asked it to write an opener in that tone, using the same five real examples as reference, in the same prompt. She ran it against a dry, deadpan bio.

It worked. Not "pretty good." It read like something the actual person in the bio would have typed.

She tried three more bios. All three landed. Total time from opening her laptop to a working opener: about twenty minutes.

Hand sketched comparison diagram titled Same request, two different prompts. Left panel, a document icon labeled Zero-shot, caption one instruction, no examples, generic opener every time. Right panel, a document icon labeled Few-shot, caption same instruction plus five real voice examples, matches on the first try.
The instruction never changed between February and June. What rode along with it did.

Meara messaged Auden that afternoon. "Hey, quick question, didn't we already try voice matched openers and it didn't work? I just got it working in twenty minutes with five examples in the prompt. Did we test it with real examples, or just describe the voice and ask for it?"

That question was the whole problem, in one sentence. Nobody had asked it in February.

We did not lose Voice Match that Thursday in February. We lost four months of treating one afternoon's test as a fact nobody had to check again.

Auden's first instinct, reading Meara's message, was to ask Rieko to just spend another sprint on it, no new instructions attached. She caught herself. That would have handed Rieko the exact same test she'd already run, and probably produced the exact same honest, wrong sounding "no." The other option that crossed her mind, waiting for whatever model upgrade the vendor announced next quarter, wouldn't have helped either. The model sitting there in February could already do this. Nobody had asked it the right way.

The team retested properly in July: five real examples of a user's own voice, pulled from that user's own past sent messages, riding along with the instruction in the same prompt. Over a two week test, Voice Match openers pulled a 34 percent reply rate, against the old generic openers' 22 percent. It shipped in August. By September, the "doesn't sound like me" complaint theme had dropped to 7 percent of critical feedback, the lowest it had been in over a year.

Hand sketched labeled parts diagram titled What the working prompt actually had in it. Center icon a document labeled The Prompt That Worked, with four labeled parts radiating around it: user's bio, match's profile, five voice examples, one instruction line.
The working prompt was not a cleverer sentence. It was the same sentence, plus four things nobody had added before.

The decision Auden would take back sits in that Thursday in February, not in Rieko's afternoon of testing. It was reading one honest report and writing it into the plan as a settled fact, with no prompt attached, no output attached, no note on what "reliably" would even need to mean. That decision made sense in the moment. Rieko was good, she'd tried it, and asking her to defend a three prompt test line by line would have felt like not trusting her.

What I would tell myself, sitting in that Thursday review: trusting Rieko and checking her test were never the same choice. I could have done both. Four months, and the twelve extra points of reply rate sitting there the whole time, is what it cost to only do the first one.

TRACE, so one afternoon doesn't quietly become a fact for four months

Not a way to prove Rieko was wrong. TRACE is what forces a receipt to sit next to every "can't," so a real test and an untested guess never look the same on paper.

TTimeline. Lay out exactly when, and what got tried at each point.
February: one afternoon, three zero-shot prompts, "can't" written into the plan Thursday. Shelved for four months, complaint theme flat near 18 percent the whole time. June: a different engineer, on an unrelated project, retests with five real examples in twenty minutes.
The gap that mattered wasn't how the model performed. It was how long a single untested claim sat there unquestioned.
RRecut. Slice "cannot" apart before trusting it as one claim.
Four different real meanings hide inside the same word: a genuine wall no prompt fixes, an approach nobody's tried yet, something the model can do but not reliably enough, or "I don't want to spend the time finding out." Rieko's report only ever supported the second one.
This is the recut that matters here: not by segment or by user, but by which of the four meanings a "can't" is actually standing in for.
Hand sketched icon list titled Four things the model can't can mean. Four numbered rows. One, a real wall no prompt fixes it. Two, this row emphasized in coral, confirmed nobody tried the right approach yet. Three, it works, just not reliably enough. Four, I don't want to find out in disguise.
Only one of the four turned out to be true here. The other three were never tested, so nobody actually knew which was which.
Complaint theme "doesn't sound like me," share of critical feedback, February to September
20% 15% 10% 5% 0 Feb Mar Apr May Jun, retest Jun Jul Aug Sep
Shelved, flat near 18%Retest monthAfter the fix shipped
Four flat months, then a straight decline once the real cause got tested and fixed. The number never moved on its own, someone had to go check it.
AAssume nothing. Don't assume "cannot" always means the same thing.
Auden's mistake wasn't trusting Rieko. It was assuming her "cannot" meant "this is architecturally impossible," when what she'd actually said, underneath the plain English, was "one plain instruction, tried three times, didn't work." Those are not the same claim, even though they used the same word.
"Cannot" gets used loosely under time pressure to cover four different real situations. It's the asker's job to find out which one, not the answerer's job to specify it unprompted.
CCause candidates. The real questions, not a list of everything possible.
Has this been tried, and how, one plain instruction or real examples in the prompt? What would "reliably enough" actually look like, and was that bar ever tested against? Is there a workaround architecture, like pulling out the hard part first and generating against it, that nobody tried?
Any one of these three questions, asked in February, would have surfaced that Rieko had only tried the cheapest version of the idea.
EEvidence test. The one question that separates a wall from a guess.
Has anyone actually run a real test against this exact case, with real examples of the target voice sitting in the prompt, or is "the model can't do voice matching" a belief drawn from how models generally behave with a plain instruction? For Kerning, the honest answer in February was: nobody had.
This is the strongest move in the whole framework. It's checkable against something real, not a guess about which explanation feels more likely.
Reply rate, generic opener versus Voice Match opener, two week test
30% 20% 10% 0 22% Generic opener 34% Voice Match opener
Old generic openerVoice Match, five examples in prompt
Twelve extra points of reply rate had been sitting untested for four months, waiting on four questions nobody asked in February.
Hand sketched metaphor scene titled Most can't moments are an untried key, not a locked door. Left panel, a plain gray box icon labeled Locked door, caption a real capability wall. Right panel, a question mark box icon labeled Untried key, caption an approach nobody tested yet.
The whole method in one picture. Before treating a "can't" like a locked door, check whether anyone has actually tried the key.

Three things worth saying plainly, since interviewers push here. Auden considered two other moves before landing on the right one. Asking Rieko to just spend another sprint trying harder, with no new instructions, would have handed her the same test again and probably produced the same honest, wrong sounding "no." Waiting for the vendor's next model release would have cost more months for nothing, since the model sitting there in February already handled the task fine once given real examples. The AI specific failure worth naming by name is this: a team's belief about what a model can do gets fixed the moment one test result gets treated as final, even though nothing about the model itself ever changed. The guardrail is unglamorous: require the actual prompt and the actual output attached to every "can't" claim before it goes into a plan as an assumption. And the trade off, accepted on purpose, is real: the few-shot version runs about three times the tokens of the zero-shot version, roughly 540 against 180, and a little slower per opener, about 1.6 seconds against 1.1. Kerning took that trade, because the alternative was a 22 percent opener that kept sounding like nobody in particular.

And if you want to be sure it really works, try it somewhere else

Same five letters, a nonprofit grant writing tool instead of a dating app, and this time the untried thing isn't examples in the prompt. It's real reference text nobody thought to attach.

Grantwell drafts grant narratives for small nonprofits, matched to a specific funder's own preferred tone, some funders want "systems change" language, others want plain "direct service" language, checked by a human reviewer before anything gets submitted. Nadeen Vaughn runs product there, and hit a version of Auden's exact mistake ten weeks into a push to make Grantwell's drafts sound like they actually understood each funder.

Hand sketched flow diagram titled Grantwell's fix wasn't examples, it was reference text. Five steps left to right: Guidelines, Retrieve excerpts, this step emphasized, Draft, Review, Submitted.
Different product, different missing ingredient. The shape of the mistake was identical.

Solweig Wrixon, the engineer who owns Grantwell's drafting model, tried matching a funder's tone with one description in the prompt: "Write this narrative in the style this funder prefers: systems focused, data led, avoids charity language." The draft came back sounding like generic nonprofit writing, the same on every funder. She tried it twice more, same result, and reported it in Friday's sync: "The model can't reliably match a specific funder's tone. I described what they want and it just writes the same way regardless." Nadeen wrote it into the roadmap as a limitation and moved the feature down the list. It sat there for ten weeks.

Mapped onto TRACE, the diagnosis ran the same shape as Kerning's, with one real difference in the answer. The timeline showed one afternoon of testing, then ten weeks of silence, then a new hire asking why nobody had tried using the funder's own past funded proposals as reference material. The recut turned up the same four meanings of "cannot," and this case also confirmed meaning two, an untried approach, but a different untried approach than Kerning's: not few-shot examples of a style, but retrieval, pulling three real excerpts from that funder's own previously funded grants into the prompt as reference text. The assumption Nadeen corrected: "can't match tone" had sounded like a hard style transfer wall, when it actually meant nobody had ever shown the model what that funder's own writing looked like. The evidence test gave the same kind of answer Kerning's did: pull real reference excerpts from that funder's past awards, put them in the prompt, and rerun the exact case that failed before. It worked on the first try.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: ask exactly what got tried, one plain instruction or real examples and reference text in the prompt, before accepting that something can't be done.
Cost: no time to trace a real incident. Ask one question instead: has anyone run a real test against this exact case, or is this a belief about how models behave in general?
The model got better, for real: say the vendor ships a bigger, sharper model next quarter. An untested "can't" from the old model doesn't automatically get retested just because a new one arrived, someone still has to go run the test.

Where people run it wrong.
They treat "I tried it once and it didn't work" as the same thing as "it's not possible."
They ask the engineer to try harder without saying what to try differently, so the exact same attempt gets repeated.
They stop trusting the model on an entire feature area once one narrow attempt fails, instead of narrowing the doubt to the one approach that actually got tested.

How to use it live. When an interviewer throws this at you cold, buy two seconds by asking one thing back: "when they say it can't, do they mean they tried it and it came back wrong, or that they're pretty sure it would?" That question alone is usually exactly what a question shaped like this one is listening for.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a question asking what to check before accepting "the model can't do that"?
Tap to flip
ANSWER
TRACE: timeline, recut, assume nothing, cause candidates, evidence test. Built for ruling out an easy sounding explanation before it becomes a fact.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Auden Colter, who owns Kerning's opener roadmap, and Rieko Cleary, the senior engineer whose honest one afternoon test got treated as final. Meara Osei-Tutu found the real fix by accident four months later.
3 · THE TIMELINE
What happened in February, and what happened in June?
Tap to flip
ANSWER
February: three zero-shot prompts in one afternoon, "can't" written into the plan, shelved. June: a different engineer tries five real voice examples in the same prompt and gets a working opener in about twenty minutes.
4 · THE RECUT
Name the four things "the model can't do that" can actually mean.
Tap to flip
ANSWER
A genuine wall no prompt fixes, an approach nobody's tried yet, something the model can do but not reliably enough, or "I don't want to spend the time finding out" wearing a capability claim's clothes.
5 · CAUSE CANDIDATES
What are the real diagnostic questions to ask before accepting "cannot"?
Tap to flip
ANSWER
Has this been tried, and how? What would "reliably enough" look like, and was that bar tested? Is there a workaround, like splitting the task into steps, that nobody tried?
6 · THE OLD DECISION
What decision would Auden take back?
Tap to flip
ANSWER
Reading Rieko's honest report and writing it into the plan as a settled fact, with no prompt attached, no output attached, and no note on what "reliably" would need to mean.
7 · THE EVIDENCE TEST
What's the one question that separates a real capability wall from an untested assumption?
Tap to flip
ANSWER
Has anyone actually run a real test against this exact case, with real examples in the prompt, or is this a belief about how models behave in general?
8 · CROSS-PRODUCT TRANSFER
Section 4 runs TRACE again on a different product. Which one, and what was the untried fix?
Tap to flip
ANSWER
Grantwell, a grant writing tool for small nonprofits, run by Nadeen Vaughn. The untried fix wasn't examples of a style, it was pulling real excerpts from a funder's own past funded grants into the prompt.

Check yourself Score: 0 / 0

Multiple choice
1. Which of the four meanings of "the model can't do that" actually applied at Kerning in February?
  • A. A genuine hard wall the model could never cross.
  • B. An approach nobody had tried yet, real voice examples in the prompt.
  • C. Something the model could do, just not reliably enough on a tested bar.
  • D. Rieko simply didn't want to spend the time on it.
Show hint
Check what Meara actually changed in June.
Show answer
B. Rieko only ever tried a plain, zero-shot instruction. The fix was real examples of the target voice sitting in the prompt, an approach nobody had tested in February.
True or false
2. True or false: Rieko was wrong to say "the model can't do this" in February.
  • True
  • False
Show hint
Look at the Assume nothing step.
Show answer
False. Her report was honest and accurate about what she tried. The mistake wasn't her test, it was Auden treating one afternoon's result as proof of a hard wall instead of one data point.
Fill in the blank
3. Rieko's three tries used ___ examples. Meara's one try used ___, and it worked in about ___ minutes.
Show hint
Look at the title of this answer.
Show answer
Zero. Five. Twenty. The only real difference between the failed test and the working one was what rode along with the same plain instruction.
Short answer, name the reversal
4. What old decision would Auden take back, and why did it make sense at the time she made it?
Show hint
Look at the key point box titled "The choice I would take back," in Let's learn.
Show answer
Model answer: Writing Rieko's honest one afternoon report into the plan as a settled fact, with no prompt or output attached. It made sense because Rieko was good and had genuinely tried it, and pushing back would have felt like distrust rather than diagnosis.
Short answer, apply it yourself
5. Think of a time someone told you an AI tool "can't" do something you wanted. What's one question from this answer you could have asked before believing it?
Show hint
Think about what was actually tried, not just what the result was.
Show answer
Model answer: Something like "did you try it with a real example of what you wanted, or just describe it and ask?" Most "can't" claims turn out to rest on the plainest possible attempt, never the strongest one.
Short answer, work the numbers
6. If Meara's prompt had used only two voice examples instead of five, would it likely have worked as well? Why or why not?
Show hint
Think about what "reliably enough" would need to mean, from the Cause candidates step.
Show answer
Probably close, though less certain. Two examples might show a voice's shape but leave more room for the model to average toward generic. This is exactly why "reliably enough" needs to be tested at the actual number used, not assumed from a similar attempt.
Before you close the answer
Why this works
Tests whether you'll write "the model can't" into a plan as settled fact, or ask for the receipt first. Most candidates jump straight to brainstorming workarounds instead of checking whether any real evidence exists at all.
Follow-up traps
"Isn't this just second-guessing your engineers?" Response: no, the fix isn't overruling Rieko, it's asking what she tried before her read gets written into the roadmap as permanent. She's the one who would have found the missing piece too, if asked the right question.

"What if you ask all these questions and it really is a hard wall?" Response: then you've spent twenty minutes and gained a documented test artifact, cheap insurance against exactly the four month freeze that happened here.
If pressed
The five voice examples that made Voice Match work weren't written by hand or picked from a style guide. They were pulled straight from each user's own past sent messages, which is exactly why they carried the right voice without anyone needing to define what "dry" or "warm" meant in the abstract.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more