ConceptIntermediateModel Fluency & the AI PM Role / The AI literacy baseline every PM needs / #8

What is a system prompt and why is it a product asset rather than a code detail?

FLIPS · a wording fix to the benefits section of Lanternway's system prompt, Hollowfield Systems' AI search assistant for Millport Freight's employees

Lanternway is Hollowfield Systems' internal search assistant: point it at a company's HR handbook, IT steps, and expense rules, and an employee gets a plain answer in seconds instead of hunting through five documents. Millport Freight, a logistics company, has run it for two years. Elspeth Ferraday owns Lanternway's system prompt, the file that decides how it's allowed to answer. Godwin Adesanya runs benefits for Millport Freight. He's the one who found out what a single wording change had been quietly telling people.

The direct answer
A system prompt is the standing instructions that shape every answer before a single employee types a word, not a detail buried in a config file. The exact same model, given two different system prompts, can be a careful assistant or a confidently wrong one, with zero code changed. Treat every wording change to it like any other product change: a second reviewer, a test against real hard questions, and a small staged group before it reaches everyone at once.
Do this, in order
  1. Treat every system-prompt wording change like a real product change: a second reviewer, a test against real hard cases, a staged rollout.Why: this is the actual reversal; skip it and whoever edits the wording last is quietly setting policy for every employee who asks tomorrow.
  2. Give the system prompt its own version history, separate from the app's code history.Why: without one, nobody can point at the exact wording that was live the day a wrong answer went out.
  3. Build a golden set of real, hard past questions, and add every production failure to it for good.Why: a test built from easy examples passes by design, which means it catches nothing that actually matters.
  4. Watch how often the assistant states a policy conclusion outright instead of deferring, on sensitive categories.Why: a jump toward confident answers is the early warning, weeks before a wrong-answer count would ever move.
  5. Roll a system-prompt change out to one small group first, not the whole company at once.Why: staging catches a bad wording change in hours, against the weeks it takes for someone to notice their pay was wrong.
  6. Leave harmless wording, greetings, formatting, non-sensitive categories, on a lighter check.Why: not every wording change carries the same risk, and treating a "hi there" tweak like a policy change wastes the review budget where it doesn't matter.

How to answer this, stage by stage

Nobody's grading whether you can define "system prompt" like a glossary entry. They're grading whether you'll notice the moment a wording change starts deciding something a lawyer would call company policy.

1
Ground it in one product and one wording change
Say it like this
"Let's make this real. Hollowfield Systems built Lanternway, an AI helper that reads a company's internal pages and answers a worker's question in plain words. Elspeth Ferraday owns its system prompt. Millport Freight is one of the companies running it. This is what happens the day a small wording fix to that prompt quietly changes what counts as company policy."
Why this works
A named product, owner, and client stop the answer from staying a vague statement about "prompt engineering."
2
Say your structure out loud
Say it like this
"I'll run this as FLIPS. Find the person whose habit changes. Locate what she stopped doing because it worked. Identify the flip, the exact verb that snaps. Pinpoint the old decision that only made sense before. Show the replay with the fix in place."
Why this works
Two sentences show the interviewer you have a plan, not just an anecdote waiting to happen.
3
Reframe what the question is really asking
Say it like this
"This isn't really asking for a dictionary definition. It's asking what a system prompt actually controls, and what happens the day someone edits it the way they'd edit a code comment, instead of the way they'd edit the one thing that decides what the product is allowed to say."
Why this works
Compresses the whole answer into one breath, before a single detail can bury it.
4
Give the one decision
Say it like this
"So here's what I'd actually do. A system prompt is the standing instructions behind every answer, so I'd treat every change to it like a real product change: a second reviewer, a run against real hard questions, and a small staged group before it reaches every employee at Millport at once."
Why this works
This is the direct answer, spoken plainly, before the story arrives to justify it.
5
Prove it with a compressed failure
Say it like this
"Here's what happens without it. Elspeth tightens the benefits wording to sound more sure, bundled into a pull request fixing something unrelated, shipped the same Thursday. Millport's leave policy has a tenure line the old wording used to defer on. The new wording doesn't reach it. Over the next five weeks, fourteen employees under ninety days get told 'yes, you qualify' for paid leave they don't have yet. Godwin Adesanya catches it because payroll flags six correction requests in one week."
Why this works
Four sentences carry a full incident that a plain retelling would take a page to earn.
6
Say what you'd watch, and what you'd leave alone
Say it like this
"Going forward, I'd watch how often Lanternway states a policy conclusion outright instead of deferring, on the sensitive categories, week over week. And I'd leave the greeting and the formatting wording alone completely. A tweak to 'hi there' doesn't need a reviewer and an eval run, it needs nothing more than what it already gets."
Why this works
Shows you're thinking past launch day, and that the fix is targeted, not blanket paranoia.
7
Close on the one line
Say it like this
"So: a system prompt isn't a code detail, it's the one place that decides what the product is allowed to say, in plain language, to everyone who asks. Treat a wording change to it like a comment, and eventually the comment decides someone's leave, wrongly, for five weeks before anyone notices."
Why this works
Leaves the room with the actual decision, not just a well-told story about fourteen employees.

Let's learn

Say a company installs an AI helper that reads every internal page it owns, the HR handbook, the IT steps, the expense rules, and answers a worker's question in plain words instead of making them hunt through five different documents.

Knowledge spark: what is a system prompt? The standing instructions a company writes for its AI helper, before any worker types a single word. It says what tone to use, what to admit it doesn't know, and when to hand a question to a person instead of answering it itself. Change the wording, and the same model can give a completely different answer to the exact same question.

Before, a worker with a benefits question dug through a policy PDF or waited on hold with HR, and a straight answer took about fifteen minutes, sometimes a full day if HR was busy. Now the helper answers in about ten seconds. And it answers with confidence, stating the policy as settled fact instead of hedging.

Hand sketched diagram titled One system prompt, every kind of question. A document icon in the center labeled Lanternway's system prompt, with four connected labels around it: Benefits questions, IT questions, Expense questions, HR policy questions.
One file. Every kind of question the company gets asked, routed through the same wording.
Here's the turn. A few more wrong answers is not what breaks this. What breaks it is a wording change nobody reviewed, quietly deciding what counts as company policy for every person who asks that question next.

At its worst, the helper spends five weeks telling employee after employee that they qualify for a benefit they don't actually have yet, because one line of the real policy sat below the sentence the new wording stopped checking.

Employees told "yes, you qualify" before checking tenure, by week
4 2 0 1 3 4, caught Wk 1 Wk 3 Wk 5
New employees told they qualify, that week
Fourteen employees total. Nobody was watching this number on purpose. Godwin found it by accident, while doing his usual month-end payroll reconciliation.
The choice I would take back Letting a wording change to the system prompt ship inside whatever code change happened to already be open, with no second reviewer and no test against a real hard question. That was fine when the prompt was three short paragraphs about tone. It stopped being fine once the same prompt was deciding real policy answers for every category the company had.

What I would leave alone: a wording change to how the helper says hello, or how it formats a list, never needs this much ceremony. Spend the review budget on the sentences that decide money, safety, or someone's job, not the ones that decide how friendly the greeting sounds.

The lesson: a system prompt looks like a string of text sitting in a file, so it's tempting to edit it the way you'd edit a comment. But it's the one place that decides, in plain language, what the product is actually allowed to say. Anything that decides that deserves the same care as the button that submits a payment.

Now here is the same thing as a story

The short version above is what you'd actually say in the room. Read this one when you want to feel exactly what a skipped review cost, sentence by sentence.

Every Thursday afternoon, Elspeth Ferraday pulled up Lanternway's support queue and read every single ticket end to end, even the ones marked low priority. She'd written Lanternway's very first system prompt herself, three short paragraphs telling it how to greet a worker, how to cite its source document, and when to say "I'm not sure, ask HR" instead of guessing. Two years in, she still knew which sentence controlled which behavior, off the top of her head.

For the first year, every wording change to that prompt got a Slack thread and a nod from a reviewing engineer before Elspeth merged it, even a change as small as softening how the helper apologized for a bad search. Some weeks it felt like overkill. Nothing had ever gone wrong from it.

Hand sketched horizontal timeline titled Elspeth's review habit, thinning out. Four milestones: Every wording change reviewed, caption a reviewer reads it, even a tone tweak. Only big changes reviewed, caption a one line tweak just goes in. Wording rides along, caption merged into whatever PR is already open. Benefits wording rewritten, this milestone emphasized, caption shipped same day, no second reviewer.
Nobody decided, on any single day, to stop reviewing wording changes. It thinned out in three quiet steps.

By the second year, only changes that opened a whole new question category got that treatment. A one-line tone tweak just went in. By the eighteen-month mark, a wording tweak rode along inside whatever pull request Elspeth already had open that day, reviewed by whoever was reviewing the unrelated code next to it, if anyone.

Then Millport Freight's contract came up for renewal, and one number on the account dashboard bothered everyone in the room: satisfaction was fine, but forty percent of benefits questions still ended with Lanternway saying "confirm with HR." Millport wanted that number down before they signed again.

So on a Thursday, Elspeth rewrote the benefits section of the prompt. The old wording said: for any eligibility question, always tell the worker to confirm with HR, and only quote the general policy text. She tightened it: state the policy's conclusion directly when the source document reads as clear, and save the HR line for anything the document itself flags as unclear. She folded the change into a pull request that also fixed how footnote numbers rendered in citations. The engineer reviewing signed off on the citation fix. It shipped that afternoon.

Millport's parental leave policy reads, in its first line, "full-time employees are eligible for sixteen weeks of paid parental leave." Three sentences later, a clause about tenure adds: employees under ninety days get unpaid leave only, under state law. The old wording had always deferred on this question entirely, so the tenure clause never mattered to what Lanternway said. The new wording read the first line as clear enough to state outright, and never got to the tenure clause at all.

Godwin Adesanya runs benefits for Millport Freight. Five weeks after the change shipped, he ran his usual month-end reconciliation and found six payroll correction requests in a single week, employees who'd started leave expecting full pay and weren't getting it. He pulled Lanternway's chat logs to see what each of them had actually been told.

Fourteen employees, over those five weeks, had asked some version of "am I eligible for parental leave" and been told, plainly, yes. Not one of them had reached ninety days.

We didn't give fourteen employees a wrong sentence. We gave several of them a wrong plan for their family's next three months.

I want to say the problem was one careless edit. It was a real cause, but it's not really the story. Elspeth never had a rule she was slowly bending. She had a switch: either a wording change went through a second reviewer and a real test, or it didn't. By month eighteen, there was no version of a quick prompt fix that still got the first one.

Here's the decision I'd take back. Not "read more carefully," and not "hire a reviewer." The actual decision got made in Lanternway's first month, when the prompt was three paragraphs and every tweak was cosmetic: reviewing a wording change felt like the same weight as reviewing a line of code, so it lived in the same pull request, checked by whoever was already looking at that PR. Nobody revisited that once the prompt grew to cover a dozen policy categories with real money and real leave attached to them.

Run the same Thursday again, with one thing changed back in month one: a wording change to the system prompt is its own pull request, tagged separately, and it can't merge until it clears Lanternway's golden set, forty real past questions the model has to answer correctly, including, now, the tenure-exception question added the week after this incident. Elspeth drafts the same confidence-tightening wording, for the same reason, under the same pressure from the same renewal meeting. The golden set catches it the same day: the test case answers "yes, you qualify" for a worker at day forty, which is wrong. She adds a tenure check to the wording instead of removing the caveat outright. It ships two days later. Zero employees get the wrong answer.

One design finds this five weeks and fourteen employees later, from a payroll spreadsheet. The other finds it the same afternoon, from a test that was built to catch exactly this.

What I'd tell my past self, the one who filed that first wording tweak next to a citation bug because splitting it into its own pull request felt like busywork: "it's just wording" is true right up until the wording is the only thing standing between a worker and a policy that isn't actually theirs yet. You don't find out which sentence that was until afterward.

FLIPS, or how a wording fix quietly becomes a policy

Not a trick to sound structured. FLIPS is what makes you notice the moment a sentence in a config file starts deciding something a lawyer would call company policy.

Hand sketched numbered list titled FLIPS, the five questions in order. Five rows: F find the person, whose habit is this. L locate the habit, what did she stop doing. I identify the flip, what verb snaps, this row in orange. P pinpoint the old decision, what only worked before. S show the replay, same day new design.
Four setup and payoff letters, and one hard question in the middle of all of them.
FFind the person. Whose habit is this?
Elspeth Ferraday, who wrote Lanternway's first system prompt and has owned it for two years at Hollowfield Systems.
The flip belongs to whoever actually has the keys to the wording, not whoever happens to be reviewing the PR it rides along in.
LLocate the habit. What did she stop doing?
She stopped routing every wording change through a second reviewer and a real test against hard questions, once eighteen months of small tweaks had shipped fine.
That habit cost nothing while the prompt was three short paragraphs. It became the thing that quietly stopped happening once the prompt covered a dozen categories that mattered.
IIdentify the flip. What verb snaps?
Old setting: a wording change gets a second reviewer and a check against real questions before it ships. New setting: a wording change rides along inside whatever code PR is already open, checked by whoever happens to be reviewing that, if anyone. Nothing in between: either the prompt is a reviewed product surface, or it's a string anyone can edit on the way to somewhere else.
This is the answer to the question in one line. A system prompt stops being a code detail exactly when someone treats reviewing it as optional, because nothing broke last time.
PPinpoint the old decision. Which choice made sense before?
Deciding, in Lanternway's first month, that a wording change carried the same review weight as a line of code, and could ship in the same pull request, checked by whoever was already there.
"Have someone glance at it sometime" would be a new dial. Giving the wording its own pull request and its own gate is the decision taken back.
SShow the replay. Same day, better ending?
Same Thursday, same renewal pressure, same wording change drafted. This time it's its own pull request, and it fails Lanternway's golden set the same day, on the exact tenure-exception case added after the first incident. Fixed and shipped two days later, with zero employees told the wrong thing.
The replay ends in a count: same day and zero, not five weeks and fourteen.
Hand sketched two panel comparison titled The I step, in one picture. Left panel a gauge icon labeled Every wording change, caption gets a second reviewer before it ships. Right panel a box icon in orange labeled Wording just rides along, caption in whatever PR is already open, no second look.
Elspeth's review habit was never a dial easing down. It was a switch, and eighteen months of quiet tweaks flipped it without anyone deciding to.
Hand sketched two panel comparison titled Same five weeks, replayed with the new design. Left panel a document icon in orange labeled Old design, caption 14 employees told yes you qualify by mistake. Right panel a document icon in green labeled New design, caption 0 employees told the wrong thing, caught before it shipped.
Same five weeks. Same kind of wording change. The only thing that moved is what stood between the wording and the employees.
Days between a risky wording change shipping and someone catching it
40 20 0 35 days Old design Same day New design
No separate review laneGolden set + staged rollout
Same kind of wording change, two very different clocks. The gap between 35 days and same day is the entire design decision.
Hand sketched full page metaphor titled What we assumed, and what was true. Left panel a gauge icon labeled What we assumed, caption she gets slowly more careful rewriting policy wording. Right panel a box icon in orange labeled What was true, caption a wording change either clears the eval or it ships blind, nothing between.
The whole answer, in one picture. Nobody designed a dial. Everybody got a switch.

Three things worth being direct about, since this is where the real judgment sits. We considered the obvious fix first: just add a disclaimer banner, "Lanternway's answers aren't guaranteed," instead of a review gate. Rejected, because a banner doesn't change what the sentence said with full confidence, the employee still walks away with the same wrong plan in their head, just with a legal footnote nobody reads. The AI-specific failure worth naming is silent behavior drift from a wording change alone: no error, no crash, no code diff a normal review would ever flag, just a different sentence coming out of the same model. The guardrail is a golden set of real, hard past questions that every wording change has to clear before it ships, not a person's read-through. And there's a real trade-off, accepted on purpose: a wording change that used to ship in about an hour now takes at minimum a day, one eval run plus a short staged window. That's slower for the tweaks that really are harmless, and we don't get to know in advance which ones those are.

And if you want to be sure it really works, try it somewhere else

Same five letters, a different kind of overload. This time the team already built a review gate, and it still didn't catch the failure, because of what the gate was actually tested against.

Kestleworth Field Services runs Ductline, an AI assistant technicians open on a tablet in the truck. Point it at a make and model and it reads the equipment manual, then walks the technician through the fix, including any high-voltage safety warning buried inside it. Ottakar Draye has kept Ductline's golden eval set, the questions every wording change has to pass, for just over a year.

Knowledge spark: what is a golden eval set? A fixed pile of real, hard questions with known right answers, that a model has to get right before a change ships. It's only ever as good as what's actually sitting inside the pile.

For the first ten months, every new case in that set came straight from a real manual: the actual page, appendix and all, pasted in whole. Building one that way took Ottakar the better part of an afternoon. Building one from memory, a clean two-paragraph version of the same scenario, took ten minutes and passed the model just as well. Nobody was watching which method he used, because both kinds of case had always caught the bugs they were supposed to catch.

A technician in the field opened Ductline on a discontinued heat-pump unit with a wiring warning tucked into a footnote on page thirty-four of a sixty-one-page manual. Ductline walked him straight past it. Nothing in the golden set had ever included a warning shaped like that, because sixty-one of Ductline's eighty cases by then were Ottakar's own tidy rewrites, and a tidy rewrite never keeps a footnote.

Hand sketched two panel comparison titled What the golden set is actually built from. Left panel a document icon labeled Old golden set, caption 61 of 80 cases are Ottakar's own tidy rewrites. Right panel a document icon in green labeled New golden set, caption every case pasted straight from the real manual, appendix and all.
Same idea, different failure shape. A review process is only as honest as what it actually tests against.

F · Ottakar Draye, who has kept Ductline's golden eval set for a year, at Kestleworth Field Services.
L · He stopped building every new golden-set case from the real source manual, once a quick, clean rewrite had passed the model just as reliably, a dozen times running.
I · The pre-editing flip, a different shape from Elspeth's. Old setting: builds each case from the real document, page numbers and appendix included. New setting: writes a short, tidy version of the scenario himself, because the real manual is a hassle to extract from. No middle: either the golden set is built from real documents, or it's quietly built from someone's summary of them.
P · When the golden set started at twelve cases, a hand-typed clean example covered the idea just as well as the real page. Nobody revisited that shortcut once the set reached eighty cases and the manuals behind it got messier and longer.
S · New rule: every golden-set case gets pasted in directly from the real source document, never rewritten by hand. The next messy manual, a different unit with a wiring note buried in a footnote, gets caught in the eval run the same week it's added to the set, not months later from a technician standing in a driveway.

What finally surfaced it A new hire, on his first week of ride-alongs, asked why Ductline hadn't mentioned the warning on page thirty-four. Ottakar checked the golden set. Sixty-one of its eighty cases turned out to be his own clean rewrites. Not one of them had a footnote.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: a review process only tests what's actually sitting in the test pile. If the pile is all clean examples, the process is theater with extra steps.
Cost: no budget for a dedicated reviewer. Pasting the real source straight into the golden set costs nothing extra, it's the same copy-paste as writing a clean version, just without the tidying.
The model got better, for real: say Ductline's overall accuracy hits 99 percent next year. Doesn't matter. A golden set built from easy examples will still miss a real footnote, no matter how good the model behind it gets.

Where people run it wrong.
They treat "we have a review process" as the finish line, without ever checking what the review actually tests against.
They let whoever's fastest at writing a test case decide how realistic that case gets to be.
They wait for a rewritten case to fail before questioning it, when a rewritten case is built, by definition, to pass.

How to use it live. When an interviewer asks what a system prompt actually is, ask yourself one thing before answering: who's allowed to change it, and does anything real stand between that person and production? If the honest answer is "nobody, really," that's the whole question, answered.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
An over-trust flip. Instead of checking every wording change, the review habit that once caught problems gets skipped entirely, because many small tweaks in a row shipped fine and nothing ever broke.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Elspeth Ferraday, product manager who wrote Lanternway's first system prompt and has owned it for two years at Hollowfield Systems.
3 · THE HABIT
What did she stop doing because it worked?
Tap to flip
ANSWER
She stopped routing every wording change through a second reviewer and a real test against hard questions, once eighteen months of small tweaks had shipped fine.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
A wording change gets a second reviewer and a real test before shipping, versus a wording change rides along inside whatever PR is already open, checked by whoever's already reviewing that, if anyone. Nothing in between.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Deciding, in Lanternway's first month, that a wording change carried the same review weight as a line of code, and could ship inside the same pull request, checked by whoever was already there.
6 · THE NUMBER
Fill in the blank: ___ employees were told "yes, you qualify" over ___ weeks before Godwin caught it. After the fix, the same kind of change is caught in ___, with ___ employees affected.
Tap to flip
ANSWER
14 employees; 5 weeks. After the fix: the same day, with 0 employees affected.
7 · THE REPLAY
Same bad day, new design, what changes?
Tap to flip
ANSWER
The same wording change is drafted for the same reason, but this time it's its own pull request and fails the golden set the same day, on the tenure-exception case added after the first incident. Fixed and shipped two days later. Zero employees told the wrong thing, instead of 14 over 5 weeks.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs FLIPS again on a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Ductline, Kestleworth Field Services' AI assistant for HVAC technicians. The pre-editing flip: Ottakar Draye stopped building golden-set test cases from real messy manuals and started writing his own tidy, simplified versions instead.

Check yourself Score: 0 / 0

Multiple choice
1. Why did Lanternway start confidently answering a leave-eligibility question it used to defer on?
  • A. The underlying model was upgraded to a newer version.
  • B. A wording change let it state a policy conclusion directly whenever the source document read as clear, and it never reached the tenure clause three sentences later.
  • C. Millport Freight changed its actual parental leave policy.
  • D. Godwin Adesanya turned off the HR-deferral setting by accident.
Show hint
Check the story, right after Elspeth rewrites the benefits wording.
Show answer
B. The wording change made "reads as clear" the bar for a direct answer, and the top-line sentence read as clear even though a later clause changed the real answer for some employees.
True or false
2. True or false: the real failure in this story is that Elspeth wrote a bad sentence.
  • True
  • False
Show hint
Look at the I step in the framework recap.
Show answer
False. The wording itself was a reasonable attempt to fix a real complaint. The failure is that nothing outside Elspeth's own read-through stood between that sentence and every employee at Millport.
Fill in the blank
3. Over five weeks, ___ employees under ninety days of tenure were told they qualified for Millport's paid parental leave. Godwin caught it after payroll flagged ___ correction requests in a single week.
Show hint
Check the two-blocks diagram and the line chart in Let's learn.
Show answer
14; 6. The gap between 5 weeks undetected and a same-day catch is the entire design decision this answer is arguing for.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look at the key point box titled "The choice I would take back," in Let's learn.
Show answer
Model answer: Deciding, in Lanternway's first month, that a system-prompt wording change carried the same review weight as an ordinary code change and could ship inside the same pull request. It made sense then because the prompt was three short paragraphs about tone, and treating every tweak as its own formal review would have been overhead for zero real risk.
Short answer, apply it yourself
5. Think of a config value, a template, or a piece of copy at your own job that anyone can edit without a second look, because it's "just text." What's the first sign it's grown into something that deserves real review?
Show hint
Look for the moment it stopped being cosmetic and started deciding an outcome someone could be wrong about.
Show answer
Model answer: A saved email template for declining a refund request. It starts as wording nobody checks. It outgrows "just text" the day someone quietly adds a sentence that promises a specific dollar amount or timeline, because now a customer can hold the company to it.
Short answer, work the number
6. If the tenure exception had been the first line of Millport's policy instead of buried three sentences later, would the same reversal, its own reviewed pull request plus a real test, still be worth building? Why or why not?
Show hint
Separate this one incident's shape from the structural risk that caused it.
Show answer
Yes, for a different reason. A clearer policy would lower the odds of this exact mistake, but it wouldn't remove the structural risk: nobody outside the author checking a wording change before it reaches everyone. The gate still has to catch whichever mistake comes next.
Before you close the answer
Why this works
Tests whether you can define "system prompt" like a glossary entry, or recognize it as the one place, in plain language, that decides what the product's allowed to say to everyone who asks. Most candidates can do the first. Fewer can say why that makes it a product surface, not a config value.
Follow-up traps
"Isn't this just extra process for a wording tweak?" Response: only if you already know in advance which tweak is harmless, and the whole point is you don't, not on a sentence that touches money, safety, or someone's job.

"Why not just have the model always defer, to be safe?" Response: that's the exact frustration that caused this. Forty percent of Millport's benefits questions were already ending in "confirm with HR," and always deferring just moves the cost from wrong answers to a helper nobody trusts enough to open.
If pressed
Lanternway's golden set doesn't just check that a final answer is correct, it checks that the model still reaches every clause a source document actually contains, not just the first sentence that reads as clear. That's the specific property the tenure-exception test case exists to catch, and it's a harder bar than "the answer happened to come out right."
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more