InterviewFoundationalModel Fluency & the AI PM Role / The AI literacy baseline every PM needs / #4
Explain hallucination in one paragraph a sales team could repeat accurately.
SPARK · the one paragraph every Corvestra rep says when a buyer asks if Fernline, their AI phone agent, ever makes things up
Corvestra builds Fernline, an AI voice agent that answers a company's customer-service phone line. Moraine Home Services dispatches technicians for plumbing, HVAC, and appliance repairs under home warranty plans, and is deciding whether to let Fernline take its calls. Before Moraine signs anything, its VP of Operations wants one honest answer to one question: does the thing ever just make something up.
The direct answer
Give every Corvestra rep one written paragraph, memorized word for word: what hallucination is, why it mostly happens on questions outside Fernline's approved knowledge base, what catches it before a customer hears it, and what a buyer should honestly expect going forward. Never let a rep improvise their own version of that answer, scarier or safer than the real one.
Do this, in order
Hand every rep the same written paragraph, word for word.Why: left to improvise, a rep always drifts to whichever extreme feels safest in the room, and both extremes cost Corvestra differently.
Make the paragraph name the real catch, not just admit the risk.Why: "it can be wrong" with nothing about what stops it sounds like a warning label, not a reason to trust the product.
Rehearse it against the exact question a buyer asks, not a softer version.Why: "does it ever make things up" is the real sentence, and a rep who only practiced a vaguer line freezes or goes off script anyway.
Keep the deep mechanism out of the live demo entirely.Why: a rep explaining how the model predicts words loses the room and answers a question nobody asked.
Check real call recordings for reps still going off script.Why: an unchecked paragraph quietly drifts back into everyone's own words within a quarter.
Give technical buyers a written deep dive doc instead of asking a rep to improvise it live.Why: keeps the live answer short while still giving a real answer to whoever wants more.
How to answer this, stage by stage
Nobody is grading whether you know what hallucination means. They are grading whether you can compress it into four sentences a stranger could repeat correctly, under pressure, in a room where the deal is on the line.
1
Scope it to one call, one buyer, one question
Say it like this
"Let's ground this in one demo. Corvestra sells Fernline, an AI agent that answers a company's customer-service line. I'll walk through the exact paragraph I'd give every rep to say when a buyer asks if it ever makes things up, using the moment an account exec, Tullis, actually got asked that, live, by Moraine Home Services."
Why this works
Keeps the answer from turning into generic advice about handling hard sales questions.
2
Say your structure out loud
Say it like this
"I'll run this as SPARK. Situation, what reps say today with no script. Payoff, the habit I want, one accurate answer every time. Anchor, the actual paragraph. Risk, what breaks if it leans too far either way. Keep out, what it deliberately leaves unsaid."
Why this works
Two seconds of structure tells the interviewer you have a plan, not just a good sentence you thought of once.
3
Reframe what the question is really about
Say it like this
"This isn't really a writing problem. It's about what happens the one time Fernline actually gets something wrong, months after the sale, and whether the buyer already saw that coming or feels lied to."
Why this works
This is the whole judgment in one line. Skip it and the rest sounds like copywriting advice.
4
Give the anchor: say the actual paragraph
Say it like this
"Here's the paragraph, word for word: 'Every AI voice agent, including ours, can say something wrong with total confidence. We call that hallucination. It happens most on questions the agent was never actually given the answer to, so instead of saying I don't know, it guesses. We cut down how often that happens by building every answer from your own approved documents and sending anything the agent isn't sure about to a live person instead of letting it guess. It still happens sometimes, on the rare question nobody trained it for, so every call gets logged and can be reviewed, and we recommend a human check on anything high stakes, like a payout or a coverage decision. It's not a reason to skip the tool. It's a reason to watch it the way you'd watch any new hire in their first month.'"
Why this works
This is the direct answer, said as real words a reader could check against the page, not a description of an answer.
5
Prove it with the failure that forced this
Say it like this
"Here's why this matters. Before this paragraph existed, a rep told a customer, Cobblewick Appliance Care, that Fernline never gets things wrong. Five months later, on a live call, it told a caller a roof leak was covered, a detail that was never loaded into its knowledge base. Cobblewick's renewal talks got rough, because the rep had promised something nobody could promise."
Why this works
One real number and one real name does more work than any abstract line about honesty.
6
Say what you'd measure after it ships
Say it like this
"I'd sample real call recordings every month and check how many reps still give their own version instead of the paragraph. Right now that number needs to be close to zero, not 'mostly fine.'"
Why this works
Shows you're thinking past the day the paragraph gets written, to whether anyone still uses it in week twelve.
7
Say what you'd deliberately leave out
Say it like this
"I wouldn't build the deep technical mechanism into this answer at all. A buyer's engineering team can ask about the model in real depth on a technical due diligence call, they get the long version there. On the sales floor, everyone gets the same four sentences."
Why this works
Shows judgment about where depth belongs, instead of trying to cram everything into one live moment.
8
Close on the one line
Say it like this
"So: one written paragraph, said the same way by every rep, that names the risk and the catch in the same breath, and leaves the deep mechanism for a written doc instead of a live guess."
Why this works
Restates the decision in a single breath, exactly what the interviewer needed before any follow-up.
Let's learn
The paragraph is four sentences long. It lives on a card taped inside the lid of Tullis Quilliam's laptop, right where he can read it without looking down.
Corvestra builds Fernline, an AI voice agent that answers a company's customer-service phone line instead of a person picking it up. For its first two years, Corvestra had three account executives, and all three had helped build Fernline themselves. When a prospect asked if it ever got things wrong, each of them answered from their own gut, and their gut was usually close enough, because they actually understood what the model could and could not see.
Three reps, three private answers to the exact same question, none of them written down anywhere.
Then the sales team grew to twenty two people. A call-recording audit that spring sampled forty demo calls and found eleven different answers to the same question, ranging from "it can just make things up randomly" to "it never gets things wrong, it's basically perfect."
Reps who gave an accurate, consistent answer when asked the question cold
Before the paragraph existedFirst audit after rollout
4 of 22 reps gave an accurate, consistent answer when asked cold before the paragraph existed. Within a week of it becoming mandatory, 21 did. The last rep was certified within the week after.
Knowledge spark: what does grounded mean for a voice agent?
Fernline doesn't know everything. It only knows what Corvestra loads in for each customer: warranty terms, service hours, pricing rules. Grounded means every answer comes from that loaded material instead of the model's own general guess. Ask it something outside that material, and the grounding runs out.
Here's the paragraph Praveena Yaxley, the AI PM who owns this, ended up writing. It's what every rep now says, word for word, when a buyer asks the question live.
“The paragraph, word for word
Every AI voice agent, including ours, can say something wrong with total confidence. We call that hallucination. It happens most on questions the agent was never actually given the answer to, so instead of saying I don't know, it guesses. We cut down how often that happens by building every answer from your own approved documents and sending anything the agent isn't sure about to a live person instead of letting it guess. It still happens sometimes, on the rare question nobody trained it for, so every call gets logged and can be reviewed, and we recommend a human check on anything high stakes, like a payout or a coverage decision. It's not a reason to skip the tool. It's a reason to watch it the way you'd watch any new hire in their first month.
Four parts, in this order, every time. Skip one and the paragraph tips into a warning or a promise.
Both edges lose. One loses the deal today. The other loses the account later.
The extra scary or extra confident answers were never really the problem. The problem was what a buyer remembered on the one day, months later, the agent actually got something wrong.
What it costs at its worst: five months after Cobblewick Appliance Care signed, a caller asked about a roof leak. Fernline told them it was covered under warranty, a detail that had never been loaded into Cobblewick's knowledge base at all. It wasn't a lie the model told on purpose. It was a guess, dressed up as a fact, because nothing stopped it from guessing. Corvestra's account manager spent eleven hours across three calls that month walking Cobblewick's ops team back from canceling, because the original rep had told them, flat out, "it never gets things wrong."
The choice I would take back
Corvestra's onboarding told new reps to "use your own judgment, just be honest" when a buyer asked about accuracy. That worked fine with three reps who had built the product and understood exactly what it could see. It stopped working the moment the team hired people who hadn't, and nobody replaced "use your own judgment" with an actual answer.
What I would leave alone: a buyer's technical due diligence call, where their own engineers ask about the model in real depth. That's exactly the room for the long version, confidence cut-off points, how often it routes to a person, what the logs actually capture. Forcing that same depth into the live sales demo would lose the room and never even reach the question that was actually asked.
The lesson: "use your own judgment" isn't really an answer. It's a bet that everyone's judgment stays as good as the three people already in the room. That bet pays off right up until the team grows past the room.
Now here is the same thing as a story
The short version above is what you actually say in the room. Read this one for why the paragraph had to exist at all, and what it cost before it did.
Tullis Quilliam can read a buyer's face before they finish their question. Three years at Corvestra, and he's closed accounts other reps had already written off, mostly by knowing exactly when to stop talking.
For a long stretch, that instinct was enough. Corvestra was small. Every rep had sat in on Fernline's build meetings, knew which customers had thin knowledge bases and which had deep ones, knew where the model was actually strong and where it was guessing. When a buyer asked a hard question, whatever a rep said off the cuff was close enough to true, because they weren't really improvising. They were reporting.
Then the team grew, fast, to twenty two. New hires read the pitch deck and repeated its confident language word for word. A few technical hires read the model card and answered so carefully they scared people. Nobody had ever actually written the real answer down, so everyone just filled the gap with whatever felt right that week.
Word got around about what happened to Cobblewick Appliance Care before anyone at Corvestra called it a pattern.
Same wrong answer from the agent, two different outcomes, depending on what the buyer was told before it happened.
Cobblewick's rep had sold the account eight months earlier, on a demo that went well, closing with a line meant to sound reassuring: "honestly, it never gets things wrong." Five months into using Fernline, a caller asked about a leaking roof. Fernline told them, with total confidence, that it was covered under warranty. Nobody had ever loaded that detail into Cobblewick's knowledge base, because it wasn't in Cobblewick's actual policy. The model didn't know that. It answered anyway.
Cobblewick's ops team didn't hear "the model made an understandable mistake." They heard "you told us it never gets things wrong, and it just cost us money." Corvestra's account manager spent eleven hours across three calls that month, walking them back from canceling outright.
We didn't lose Cobblewick's trust the day the agent got the warranty detail wrong. We lost it five months earlier, in a demo, when a rep promised it never would.
The old decision sat in a much smaller meeting, back when Corvestra wrote its first onboarding guide. Someone asked what a rep should say if a buyer pushed on accuracy. The answer the room landed on was "use your own judgment, just be honest." It made sense that day. Every rep in the building actually knew the product cold. Writing a rigid script felt like it would make the team sound like they were reading from a card instead of actually understanding what they sold.
Cobblewick's incident is what finally moved this from a nice idea to a required one.
Praveena Yaxley wrote the paragraph in the weeks after Cobblewick's renewal call, tested it against the actual audit recordings, the real questions buyers had asked, not a made-up FAQ. It went from optional reading to a required, certified line in onboarding. Every rep had to say it back, unscripted, to a manager, before their first live demo.
Run the same afternoon again, months later, with the paragraph in place. Tullis is mid-demo with Sarolta Grielle, VP of Operations at Moraine Home Services, walking her ops team through how Fernline handles a warranty claim call. Partway through, Sarolta leans forward. "Does it ever just make things up?"
Tullis doesn't pause. He says the paragraph, the same four sentences Praveena wrote, the same words he said back to his manager in training. Sarolta writes one line in her notebook. By 10:14, four minutes after she asked, they're back on pricing.
What I'd tell my past self, sitting in that first onboarding meeting: "use your own judgment" isn't a real answer. It's a wish that everyone's judgment stays as good as the three people already in the room, and that wish runs out exactly when the room gets bigger.
SPARK, in one screen
Not a script to sound rehearsed. SPARK is what forces you to name the one paragraph a whole team owns, instead of trusting twenty two different people to each land on something close enough.
SSituation. Who is this person, and how does the moment go today, without a script?
Tullis Quilliam is mid-demo with Sarolta Grielle, VP of Operations at Moraine Home Services, a company that dispatches technicians for plumbing, HVAC, and appliance repairs under home warranty plans. She wants one honest answer, before she signs anything, to whether Fernline ever makes things up.
One person, one call, one real question. Never a segment of "buyers."
PPayoff. What habit do I want this to build?
I want every rep to stop reaching for their own gut and start repeating one accurate paragraph, the same words, every single time a buyer asks this. The habit is the product. A calmer sales call is downstream of that habit, not the goal itself.
Name the thing they'll stop doing. That's the payoff, not the paragraph's word count.
AAnchor. The one decision everything else hangs on.
The paragraph above, exactly as written, said by every rep, no local edits. It names what hallucination is, why it happens, what catches it, and what to honestly expect. The catch itself carries a real trade: routing an unsure answer to a live person costs a customer real minutes and staffing, versus letting Fernline guess and answer instantly. Corvestra recommends taking the slower, costlier path anyway, on anything that touches money or coverage.
Concrete enough to argue with. This is the answer to the question.
RRisk. What breaks the first time it's wrong?
Too scary and a buyer hears "unreliable" and walks, the way a prospect might after "it can just make things up randomly." Too reassuring and a real incident later feels like a lie, the way Cobblewick's did. Both cost Corvestra, just on different days.
Not "accuracy drops." What the buyer does next, in either direction.
Same underlying facts, two different rooms. The live demo gets four sentences. The technical review gets the rest.
KKeep out. What I deliberately will not build into this answer.
The deep mechanism, how the model actually predicts an answer, stays out of the live paragraph entirely. It lives in a written technical deep dive doc for a buyer's engineering team, on their own due diligence call. Reps never need it live, and reaching for it live only slows the room down and answers a question nobody asked.
Shows judgment instead of a wish to cover everything. Ties straight back to Risk: depth in the wrong room is its own kind of wrong answer.
The recap, one line per letter: situation is a buyer asking a hard question live, payoff is every rep repeating one accurate paragraph instead of guessing, anchor is the paragraph itself, tuned against real audit recordings rather than a made-up FAQ, risk is either extreme costing Corvestra on a different day, and keep out draws the line at the room, not the honesty, the deep mechanism goes in a written doc, never in a live guess.
One option Praveena rejected on the way to this design: a longer, fully technical paragraph, six or seven sentences covering confidence cut-off points and routing logic in real detail, taught to every rep as the standard line. She turned it down because a rep under pressure compresses six sentences into whichever extreme feels safe, which recreates the exact problem the paragraph was built to fix, just with more words in it. She also rejected leaving it open, "say whatever feels right, just don't lie," because that was the literal status quo that produced eleven different answers across forty calls in the first place.
New deals lost per week, citing "AI reliability concerns" in the loss note
Before the paragraph was mandatoryAfter
Four deals in one quarter were lost with "AI reliability concerns" written into the loss note, every one of them after a call where the rep had used some version of the scary line. None since the paragraph became mandatory in week seven.
And if you want to be sure it really works, try it somewhere else
Same five letters, a clinic's front desk instead of a warranty dispatch line, and this time the wrong guess isn't a roof leak, it's a patient's copay.
Palterra builds an AI phone agent that reschedules and confirms appointments for Rivenwood Clinics, a network of family medical clinics. Manoela Ulyanova hit the same wall eight months into selling it: a front office director asked, mid-demo, "will it ever tell a patient the wrong copay?" and three different reps in the room had three different answers.
Same shape of problem, a different wrong guess. Scripted holds. Improvised drifts.
Mapped onto SPARK: the situation is a front office director asking about copay accuracy, live, at one clinic network. The payoff is every Palterra rep repeating one paragraph instead of guessing. The anchor is a paragraph naming the real risk, a wrong estimate on a plan detail outside that clinic's own fee schedule, the real catch, every answer grounded in the clinic's actual fee schedule with anything uncertain routed to a live scheduler, and the honest expectation, rare, logged, reviewable. The risk is the same shape as Corvestra's: a wrong dollar figure instead of a wrong coverage line, but the same two ways to lose, scaring a director off or setting up a patient complaint later. And keep out draws the same line: the deep mechanism stays for the clinic's own IT security review, never the sales floor.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it, say the paragraph, its four parts, and what it protects against.
Cost: no enablement budget this quarter. Have whoever already owns the paragraph run it as a five minute training in the next team meeting and require every rep to say it back once, free, this month.
The model got better, for real: say Fernline's accuracy on grounded answers doubles overnight. The paragraph still matters, because a smarter model that still occasionally guesses on a detail nobody gave it can still cost a customer real money once. Reliability and model quality are two different jobs.
Where people run it wrong.
They hand a buyer the full technical FAQ as their live answer and bury the actual point under detail nobody asked for.
They write the paragraph once and never check it against real call recordings, so it quietly drifts back into everyone's own words within a quarter.
They read it flat, like a legal disclaimer, which lands worse than either extreme it was built to replace.
How to use it live. Before answering a hallucination question cold in an interview, ask yourself one thing: what's the one sentence I'd want every seller at this company saying the same way, not what's the most complete technical answer I personally could give. Naming that sentence first is usually the exact distinction a SPARK question is listening for.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a question asking you to design one thing a whole team should say, before it exists?
Tap to flip
ANSWER
SPARK: situation, payoff, anchor, risk, keep out. Built to run forward from a real gap in the room, instead of working backward from a failure.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Praveena Yaxley, the AI PM at Corvestra who writes the paragraph. Tullis Quilliam, the account exec who says it live. Sarolta Grielle, VP of Operations at Moraine Home Services, who asks the question.
3 · THE PAYOFF
What habit does the one paragraph exist to build?
Tap to flip
ANSWER
Every rep stops answering from their own gut and starts repeating one accurate paragraph, the same words, every single time a buyer asks if Fernline ever makes things up.
4 · THE ANCHOR
What's the one concrete thing Praveena owns in this answer?
Tap to flip
ANSWER
The four-part paragraph itself: what hallucination is, why it happens, what catches it, and what a buyer should honestly expect. Not any single rep's own phrasing of it.
5 · THE OLD DECISION
What decision would Praveena take back?
Tap to flip
ANSWER
Corvestra's onboarding line telling new reps to "use your own judgment, just be honest." It worked with three reps who had built the product themselves. It broke once the team hired people who hadn't.
6 · THE NUMBER
Fill in the blank: a call-recording audit sampled ___ demo calls and found ___ different versions of the same answer.
Tap to flip
ANSWER
40 calls, 11 different versions, ranging from "it can just make things up randomly" to "it never gets things wrong."
7 · THE RISK, SURVIVED
What breaks if the paragraph leans too far either way, and how does the design survive that?
Tap to flip
ANSWER
Too scary loses the deal on the spot. Too reassuring sets up a real incident later, the way Cobblewick's did. It survives because the paragraph names the risk and the catch in the same breath, so it can't quietly slide into either extreme.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs SPARK again on a different product. Which one, and what's the equivalent anchor?
Tap to flip
ANSWER
Palterra, an AI phone agent that reschedules patient appointments for Rivenwood Clinics. The equivalent anchor is a paragraph about wrong copay estimates instead of wrong warranty coverage, grounded in the clinic's real fee schedule.
Check yourself Score: 0 / 0
Multiple choice
1. Why does "it can just make things up randomly" actually cost Corvestra a deal, even though it sounds closer to honest than "it never gets things wrong"?
A. Because it's factually wrong about what hallucination is.
B. Because a buyer hears "randomly" and assumes total unpredictability, when the real pattern is narrower and catchable.
C. Because reps aren't allowed to say the word hallucination out loud.
D. Because Fernline never actually gets anything wrong.
Show hint
Look at what the paragraph actually says instead of "randomly": grounded in approved documents, unsure answers routed to a person.
Show answer
B. "Randomly" implies no pattern at all and nothing catching it. The real story is narrower and has a catch, grounding plus routing plus logs, which is exactly what "randomly" throws away.
True or false
2. True or false: the real mistake at Corvestra was letting three early reps each answer this question in their own words.
True
False
Show hint
Check "The choice I would take back" in Let's learn, and compare it to what changed between three reps and twenty two.
Show answer
False. With three reps who had built the product themselves, "use your own judgment" was a fair bet, their judgment was genuinely reliable. It stopped being fair the moment the team hired people who hadn't built anything and had never been given a real answer to fall back on.
Fill in the blank
3. Fill in the blank: a call-recording audit sampled ___ demo calls and found ___ different versions of the same answer.
Show hint
Check the bar chart in Let's learn, and flashcard 6.
Show answer
40. 11. The versions ranged from "it can just make things up randomly" to "it never gets things wrong, it's basically perfect."
Short answer, name the reversal
4. What old decision would Praveena take back, and why did it make sense when it was first made?
Show hint
Look at the key point box titled "The choice I would take back," in Let's learn.
Show answer
Model answer: Corvestra's onboarding told new reps to "use your own judgment, just be honest" when asked about accuracy. It made sense with three reps who had built Fernline and understood exactly what it could and couldn't see. It stopped making sense once the team grew to twenty two and most new hires had never built anything, so their "own judgment" was really a guess.
Short answer, apply it yourself
5. Think of a product or service you've bought where a salesperson had to answer a hard, honest question live. What's one sentence you wish every seller there had been required to say the same way?
Show hint
Look for the moment a seller either scared you off or promised more than they should have, and think about what the accurate middle sentence would have been.
Show answer
Model answer: A car dealer's line on a used vehicle's history could be one required sentence: "This car passed our inspection on these specific points, and here's the one thing we couldn't fully verify," instead of either "it's basically perfect" or a vague shrug about used cars always being a risk.
Short answer, work the number
6. If Corvestra had 60 account executives instead of 22, would one written paragraph still be the right fix, or would something else be needed too? Why?
Show hint
Look at how the problem got worse specifically because the team grew, not because the model changed.
Show answer
Model answer: Yes, still the right first move, an unwritten answer scales worse the more people say it, not better. At 60 reps you'd likely also want it built into a required certification quiz rather than checked only in a monthly recording audit, so drift gets caught automatically instead of by sampling.
Before you close the answer
Why this works
Tests whether you treat "explain a technical risk to a sales team" as a design problem, not a writing problem. Most candidates write one good paragraph and stop. The stronger answer designs for what happens in the room when a buyer doesn't like either extreme, and for what happens months later when the model actually gets something wrong.
Follow-up traps
"What if a buyer pushes past the paragraph and wants the real technical explanation, right there in the room?" Response: give it in one honest line, then bridge out: "That's a fair question, and I'd rather give you the full written version than simplify it wrong live," and send the deep dive doc same day.
"Isn't the 'new hire' line a little soft for something that could cost a customer real money?" Response: no, because the paragraph doesn't stop at the analogy. It names what actually catches the mistake before a customer feels it, grounding, routing, logging, and calls out a human check on anything high stakes by name.
If pressed
The routing cut-off isn't fixed across customers. Corvestra tunes it per how complete that customer's own knowledge base is, so a thin knowledge base routes to a person more often, which costs that customer more staff time up front and drops as their knowledge base fills in over the first few months.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.