An engineer says the model cannot do that. What questions do you ask before accepting it?
Kerning drafts opening lines and reviews message drafts for people navigating dating apps, and Auden Colter owns the roadmap for how personal those openers actually feel. This is the four months a promising idea sat shelved because one honest, one afternoon test got written into the plan as the final word.
- Ask exactly what got tried before writing "can't" into the plan.Why: a single zero-shot instruction and a real test with target-voice examples are different pieces of evidence, and only one of them settles anything.
- Ask what "reliably enough" would mean, and whether that bar got tested.Why: "the model can't do X" often really means "it missed the bar once," a smaller claim wearing a bigger one's clothes.
- Ask whether a workaround, like splitting the task into steps, got tried.Why: some tasks fail in one shot and land the moment the hard part gets pulled out and generated against on its own.
- Ask for the actual prompt and output as evidence, not a verbal summary.Why: a claim with no test attached is a belief wearing a fact's clothes, and it reads exactly like a real wall eighteen months later.
- Log every "can't" claim with exactly what was tried.Why: without a record, nobody can tell a settled wall from an untested guess when the question comes back around.
- Don't reopen a claim on a hunch, only on a genuinely new approach worth testing.Why: real capability walls exist too, and re-checking a well-tested no wastes the same time this whole method is meant to save.
How to answer this, stage by stage
Nobody is grading whether you know what a token is. They're grading whether you'll let one afternoon's test decide a roadmap, or ask for the receipt first.
Let's learn
Kerning reads a person's dating profile and a match's profile, then drafts an opening message and checks message drafts before they get sent, so starting a chat stops feeling like homework.
For its first two years, every opener Kerning wrote followed one recipe: find a topic in the match's profile, mention it, ask a light question. It worked well enough. Reply rates on those openers held around 22 percent, in a two week test cohort, close to what people got writing their own generic openers by hand. Users kept using it. But the single most repeated line in Kerning's own feedback box was some version of "it works, but it doesn't sound like me."
A product lead pitched a fix: openers written in the user's own voice, not just about the user's own topic. A senior engineer tried it, one plain instruction, no examples, three different times. Each time the model handed back the same kind of line it always wrote: pleasant, topical, generic. So the engineer reported back: the model can't do this. It just defaults to a generic compliment no matter how you ask.
Here is the turn. That report was true, and it was also not proof of anything close to what it got used for.
The idea got written into the plan as "not feasible: a model limitation" and moved to the bottom of the backlog. Nobody attached the actual prompts that were tried. Nobody wrote down what "sounds like someone's voice" would even need to look like to count as working. The sentence "the model can't do voice matching" just started getting repeated in planning meetings, the way a fact gets repeated once nobody remembers it was ever a guess.
What it costs at its worst: four straight months where the single most common complaint kept showing up in almost one out of every five pieces of critical feedback, a live, named, understood problem, sitting completely untouched, because everyone believed it had already been ruled out.
What I would leave alone: Kerning genuinely cannot know whether a match is going to reply. That's not a prompting gap. No number of examples fixes a question about someone else's future choice, so that kind of "can't" doesn't need four sharp questions, it needs to be accepted and designed around instead.
The lesson: a "no" backed by one afternoon and a "no" backed by three genuinely different approaches use the exact same word, but they are not the same kind of no. My job isn't to doubt every engineer's read. It's to make sure "can't" always comes with a receipt before it becomes a fact everyone repeats.
Now here is the same thing as a story
Say the short version out loud in an interview. Read this one when you want to feel exactly how one true sentence, said once, quietly became a fact nobody checked again for four months.
Rieko Cleary can read a bad model output and tell you, inside one sentence, whether the problem is the prompt or the model itself. Most engineers guess. She checks. In three years on Kerning's small AI team, she's the person people go to when a feature isn't behaving and nobody knows why.
Every Tuesday at ten, the product team crowds into the same glass conference room to walk the roadmap, and by the second hour the air always turns thick. In February, Auden Colter, who owns how personal Kerning's openers actually feel, brought a new idea into that room: openers written in a user's own voice, not just about the user's own topic. Someone dry and self deprecating in their bio should get an opener that sounds dry and self deprecating. Someone warm and full of exclamation points should get one that sounds warm.
Rieko liked the idea and said she'd try it that week.
She spent one afternoon on it. Three prompts, each a version of the same plain instruction: "Write an opening line in this person's voice, based on their bio, that references something specific about their match." Three times, the model handed back a pleasant, topical line that could have come from anyone's bio. Nothing dry. Nothing warm. Just generic, the way Kerning's openers had always been generic.
She wrote it up honestly, the way she writes up everything. "Tried it three different ways this afternoon. The model just defaults to a generic compliment, no matter what's in the bio. I don't think it can actually pick up on someone's specific voice from a paragraph of text."
Auden read that Thursday, before the next roadmap review. It sounded right. It sounded like the kind of thing someone who'd actually tried it would say. Voice Match got marked "not feasible: a model limitation" and moved to the bottom of the icebox, and the room moved on to the next item on the list.
Nobody wrote down the three prompts Rieko had actually tried. Nobody wrote down what "sounds like someone's voice" would need to look like to count as working. The sentence "the model can't do voice matching" just started getting repeated in planning meetings, the way a fact gets repeated once nobody remembers it was ever a guess.
For four straight months, nothing about it changed. Kerning's own feedback tool kept tagging the same complaint theme, some version of "it works, but it doesn't sound like me," and it held between 17 and 19 percent of all critical feedback every single month, February through May. Auden saw the number every month. It never occurred to her to connect it back to a feature the team had already ruled out.
In June, Meara Osei-Tutu, who'd joined the AI team two months earlier, was building something unrelated: a tone checker meant to flag when a user's draft reply didn't match the tone of the conversation so far. To teach it what "dry," "warm," and "flirty" actually looked like, she fed the model five real labeled examples of each, pulled straight from real message threads, sitting right there in the prompt.
Almost by accident, she flipped one of her own test prompts around. Instead of asking the model to label a message's tone, she asked it to write an opener in that tone, using the same five real examples as reference, in the same prompt. She ran it against a dry, deadpan bio.
It worked. Not "pretty good." It read like something the actual person in the bio would have typed.
She tried three more bios. All three landed. Total time from opening her laptop to a working opener: about twenty minutes.
Meara messaged Auden that afternoon. "Hey, quick question, didn't we already try voice matched openers and it didn't work? I just got it working in twenty minutes with five examples in the prompt. Did we test it with real examples, or just describe the voice and ask for it?"
That question was the whole problem, in one sentence. Nobody had asked it in February.
Auden's first instinct, reading Meara's message, was to ask Rieko to just spend another sprint on it, no new instructions attached. She caught herself. That would have handed Rieko the exact same test she'd already run, and probably produced the exact same honest, wrong sounding "no." The other option that crossed her mind, waiting for whatever model upgrade the vendor announced next quarter, wouldn't have helped either. The model sitting there in February could already do this. Nobody had asked it the right way.
The team retested properly in July: five real examples of a user's own voice, pulled from that user's own past sent messages, riding along with the instruction in the same prompt. Over a two week test, Voice Match openers pulled a 34 percent reply rate, against the old generic openers' 22 percent. It shipped in August. By September, the "doesn't sound like me" complaint theme had dropped to 7 percent of critical feedback, the lowest it had been in over a year.
The decision Auden would take back sits in that Thursday in February, not in Rieko's afternoon of testing. It was reading one honest report and writing it into the plan as a settled fact, with no prompt attached, no output attached, no note on what "reliably" would even need to mean. That decision made sense in the moment. Rieko was good, she'd tried it, and asking her to defend a three prompt test line by line would have felt like not trusting her.
What I would tell myself, sitting in that Thursday review: trusting Rieko and checking her test were never the same choice. I could have done both. Four months, and the twelve extra points of reply rate sitting there the whole time, is what it cost to only do the first one.
TRACE, so one afternoon doesn't quietly become a fact for four months
Not a way to prove Rieko was wrong. TRACE is what forces a receipt to sit next to every "can't," so a real test and an untested guess never look the same on paper.
Three things worth saying plainly, since interviewers push here. Auden considered two other moves before landing on the right one. Asking Rieko to just spend another sprint trying harder, with no new instructions, would have handed her the same test again and probably produced the same honest, wrong sounding "no." Waiting for the vendor's next model release would have cost more months for nothing, since the model sitting there in February already handled the task fine once given real examples. The AI specific failure worth naming by name is this: a team's belief about what a model can do gets fixed the moment one test result gets treated as final, even though nothing about the model itself ever changed. The guardrail is unglamorous: require the actual prompt and the actual output attached to every "can't" claim before it goes into a plan as an assumption. And the trade off, accepted on purpose, is real: the few-shot version runs about three times the tokens of the zero-shot version, roughly 540 against 180, and a little slower per opener, about 1.6 seconds against 1.1. Kerning took that trade, because the alternative was a 22 percent opener that kept sounding like nobody in particular.
And if you want to be sure it really works, try it somewhere else
Same five letters, a nonprofit grant writing tool instead of a dating app, and this time the untried thing isn't examples in the prompt. It's real reference text nobody thought to attach.
Grantwell drafts grant narratives for small nonprofits, matched to a specific funder's own preferred tone, some funders want "systems change" language, others want plain "direct service" language, checked by a human reviewer before anything gets submitted. Nadeen Vaughn runs product there, and hit a version of Auden's exact mistake ten weeks into a push to make Grantwell's drafts sound like they actually understood each funder.
Solweig Wrixon, the engineer who owns Grantwell's drafting model, tried matching a funder's tone with one description in the prompt: "Write this narrative in the style this funder prefers: systems focused, data led, avoids charity language." The draft came back sounding like generic nonprofit writing, the same on every funder. She tried it twice more, same result, and reported it in Friday's sync: "The model can't reliably match a specific funder's tone. I described what they want and it just writes the same way regardless." Nadeen wrote it into the roadmap as a limitation and moved the feature down the list. It sat there for ten weeks.
Mapped onto TRACE, the diagnosis ran the same shape as Kerning's, with one real difference in the answer. The timeline showed one afternoon of testing, then ten weeks of silence, then a new hire asking why nobody had tried using the funder's own past funded proposals as reference material. The recut turned up the same four meanings of "cannot," and this case also confirmed meaning two, an untried approach, but a different untried approach than Kerning's: not few-shot examples of a style, but retrieval, pulling three real excerpts from that funder's own previously funded grants into the prompt as reference text. The assumption Nadeen corrected: "can't match tone" had sounded like a hard style transfer wall, when it actually meant nobody had ever shown the model what that funder's own writing looked like. The evidence test gave the same kind of answer Kerning's did: pull real reference excerpts from that funder's past awards, put them in the prompt, and rerun the exact case that failed before. It worked on the first try.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: ask exactly what got tried, one plain instruction or real examples and reference text in the prompt, before accepting that something can't be done.
Cost: no time to trace a real incident. Ask one question instead: has anyone run a real test against this exact case, or is this a belief about how models behave in general?
The model got better, for real: say the vendor ships a bigger, sharper model next quarter. An untested "can't" from the old model doesn't automatically get retested just because a new one arrived, someone still has to go run the test.
Where people run it wrong.
They treat "I tried it once and it didn't work" as the same thing as "it's not possible."
They ask the engineer to try harder without saying what to try differently, so the exact same attempt gets repeated.
They stop trusting the model on an entire feature area once one narrow attempt fails, instead of narrowing the doubt to the one approach that actually got tested.
How to use it live. When an interviewer throws this at you cold, buy two seconds by asking one thing back: "when they say it can't, do they mean they tried it and it came back wrong, or that they're pretty sure it would?" That question alone is usually exactly what a question shaped like this one is listening for.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if you ask all these questions and it really is a hard wall?" Response: then you've spent twenty minutes and gained a documented test artifact, cheap insurance against exactly the four month freeze that happened here.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Working with ML engineers and researchers
- #1 How do you write a requirement for a team whose output is a probability distribution?
- #3 Describe how you would run a planning session when effort estimates are genuinely unknowable.
- #4 What does a healthy PM-to-research relationship look like when research timelines are open-ended?
- #5 How do you keep a research team connected to user problems without constraining their exploration?
- #6 Your ML team wants three months to improve accuracy by two points. How do you evaluate that ask?
- #7 Explain how you would run a bug triage meeting where half the bugs are model behaviour, not code.