ConceptAdvancedEval-Driven Specification / Acceptance criteria for non-deterministic output / #12

Describe criteria for consistency: should identical inputs produce identical outputs?

The direct answer
Yes, but only for the storm-warning card, the short line Cloudline drafts for a live weather warning and runs on the TV crawl, the radio script, the app push, and the website banner at the same time. Generate that one line once, at zero randomness, and push the exact same words to every channel. Leave the daily "sunny and 82" forecast blurb free to vary, because a slightly different sentence about tomorrow costs nobody anything, and a warning that quietly disagrees with itself costs a viewer the one fact they actually needed.
How to decide where identical output actually matters, in order
  1. Lock the storm-warning card to one deterministic call, reused word for word on every channel.Why: the same warning, read on two screens, has to say the same thing, or the two screens start arguing with each other.
  2. Keep the daily forecast blurb on its own separate, creative call.Why: nobody compares "sunny and 82" against yesterday's wording, so forcing it to repeat verbatim only makes the product feel broken.
  3. Say who feels each kind of miss in real minutes and calls, not adjectives.Why: "46 calls in 20 minutes" moves a fix up the list. "That felt off" doesn't move anything.
  4. Track how many times the locked warning text can repeat before people stop opening it.Why: pure sameness has its own failure mode, and only a real number tells you when you've hit it.
  5. Recheck the call once a full storm season has run, not just after launch.Why: a rule built on one bad Thursday needs to survive an actual season before anyone should trust it.

The four moves for deciding when sameness is the point

A yes-or-no consistency question wants a pick, not a rulebook for every line Cloudline writes, so PICK does the work here.

1
Ground it in one real product before naming the framework
Say it like this
"Let's put this on Cloudline, a tool that turns a weather feed into the forecast copy running on a TV crawl, a radio script, an app push, and a website banner. I want to talk about one piece of that copy specifically: the line that runs during a live weather warning."
Why this works
Stops the answer floating at "should AI ever repeat itself" and gives the interviewer one concrete piece to push on.
2
Preview the four moves before making any of them
Say it like this
"I'm going to pick a position on one specific piece first, then say who feels each kind of miss and in what units, then name which kind is actually worse, then say what would change my mind. Four moves, in that order."
Why this works
Signals a method instead of a ramble, and tells the interviewer what's coming before you start.
3
Say what the question is actually testing
Say it like this
"This isn't really 'should AI be consistent.' It's 'do you know that some lines get read once and forgotten, and some get compared, screen to screen, by someone standing in a driveway deciding if it's safe to go outside.' So I'm not going to answer for the whole product. I'm going to answer for the one piece where that comparison actually happens."
Why this works
Shows the interviewer you see past the surface ask to the real judgment being tested.
4
Give the position, with the actual mechanism in it
Say it like this
"My position: yes, for the storm-warning card, and only for that. Right now Cloudline calls the model separately for each of the four channels, so the same warning can come back worded four slightly different ways. I'd call it once, at zero randomness, and push that exact text to all four. The daily blurb keeps its own separate, more creative call."
Why this works
PICK rewards committing to a real mechanism, not a vague promise to "be more consistent."
5
Name who feels each kind of miss, then prove the asymmetry with a real case
Say it like this
"Here's the split. On a calm day, if the app says 'warm and clear' and the TV says 'sunny and mild' about the same forecast, nobody notices, nobody's plans change. But during a real tornado warning, Cloudline once had the TV crawl say the warning ends at 4:00 and the app say 4:15, same warning, same input, two separate calls to the model. A new producer spotted it and asked which one was right. In the twenty minutes it took to sort out, the newsroom logged 46 calls, and Southridge read an on-air correction eight minutes later."
Why this works
The real numbers and the real case make the asymmetry checkable, not just asserted.
6
Say what you'd leave alone
Say it like this
"I wouldn't touch the daily blurb. 'Sunny and 82' can read a little different every morning and it costs nothing. Locking that down too would just make the forecast sound like a recording, and that's a worse product for no safety gain at all."
Why this works
Shows judgment instead of blanket caution, the check most answers skip when they hear the word "consistency."
7
Name the kill criteria and close on the one line
Say it like this
"I'd revisit this if the locked warning text started repeating so often during a multi-day storm that people stopped opening it. That's a real failure mode too. But until I see that in the numbers, I'll take a boring, repeated warning over one that quietly disagrees with itself on the worst afternoon of somebody's week."
Why this works
Shows confidence without stubbornness, and ends on the line the interviewer should walk away remembering.

One more thing before the walkthrough moves on: this pick covers one piece of copy, not the whole product. Most candidates hear "make it consistent" and answer by locking everything down. Say which piece earns the deterministic pass and which ones stay free to vary, and you've shown real judgment instead of reciting a safety word.

Let's learn

Cloudline is a small box behind Southridge Broadcasting's weather segment: feed it the raw storm data, and it writes the plain-English line that runs on the TV crawl, the radio script, the app push, and the website banner.

Before Cloudline, a producer typed that line by hand for each of the four channels, checking each one against the last to make sure the wording matched, about nine minutes a warning, longer once a real storm was already moving through.

Knowledge spark: what is sampling temperature? A setting that decides how much a model is allowed to vary its own wording when it answers the same question twice. Set it to zero and it writes the exact same words every time, which is called deterministic. Turn it up and it takes more chances with phrasing, useful for a daily blurb, risky for anything that gets compared side by side.

With Cloudline, a producer gets all four drafts back in under twenty seconds. Every piece of copy it wrote, warning card included, ran through the model at the same setting, one built to keep the daily forecast from sounding like a recording: a sampling temperature of 0.7.

Here is the turn. That setting made the daily blurb sound fresh, and it worked fine there for years. But it also meant every warning card got drafted four separate times, once per channel, and each of those four calls was free to phrase the same fact a little differently. Nobody had ever asked whether the crawl and the app owed each other the same words.

We didn't ask Cloudline to disagree with itself. We just never told it not to.

At its worst, that looks like this: a tri-county severe thunderstorm warning goes out, and the TV crawl rounds the expiration time to 4:00 while the app states it exactly, 4:15, both drawn from the same National Weather Service alert. A viewer near the county line, watching the crawl, thinks the storm has already cleared by 4:00 and steps outside fifteen minutes early.

The choice I would take back We ran every piece of Cloudline's copy through one shared setting, tuned for the daily blurb, because that was the piece everyone thought about first. I would take that back. Before any warning ships, I'd pull the warning card onto its own call, locked to zero randomness, generated once and reused word for word across all four channels.
Cost, in newsroom minutes, of a wording miss
2 min 20 min +46 calls A stale blurb, caught in review A mismatched warning, live on air
The green bar is small on purpose: a daily blurb that sounds a little repetitive gets noticed and tweaked inside a normal shift, about 2 minutes, nobody outside the newsroom sees it. The red bar is what one mismatched storm-warning card actually cost: 20 newsroom minutes sorting out which time was right, 46 calls from viewers in that same window, and an on-air correction eight minutes after the mismatch went up. Neither number counts the fifteen minutes a driveway-side viewer might have acted on the wrong end time.

What I would leave alone. The daily forecast blurb, and anything else nobody reads two copies of side by side: internal drafts, archived past forecasts, and the seven-day outlook graphic, which nobody screenshots against a second channel to check if it matches.

The lesson. We built Cloudline to answer "does this sound like good forecast copy." It does. We never asked whether the same input, asked twice, owed anyone the same answer, and only the second question was the one that mattered on the one Thursday it counted.

The county line at 4:02

You don't need this to answer the question. Read it if you want to feel why the locked pass has to sit on the warning card before the next storm, not after.

Every storm season, Rutendo Chivayo checks the same three tabs from her desk at Southridge Broadcasting: the crawl, the app, and the radio script, one for each channel Cloudline feeds.

She's run product for Cloudline for five years, since before it wrote a single word on its own. Before that, a producer typed every warning by hand, once per channel, comparing each draft against the last so nothing shipped out of step. It took about nine minutes a warning, and there wasn't always nine minutes to spare.

For the first year Cloudline ran, Rutendo read all three channels side by side before anything went out, the way the old process used to force a producer to do. They always matched close enough. Routine watches, ordinary rain, nothing ever looked wrong.

So she stopped reading all three. She started skimming just the crawl, since that was the one most viewers actually saw.

Then she stopped skimming that too. If the pass check on the dashboard said CLEAR, she moved on to the next thing on her list.

Then, on a Thursday afternoon in April, a line of severe thunderstorms rolled toward three counties, one of them the county Southridge broadcasts from. Cloudline drafted the warning card the way it always did, once for the crawl, once for the app, each a separate call to the same model.

Two unequal boxes: a small calm box shows a person at home glancing at a phone and TV that word tomorrow's sunny forecast a little differently, nobody checks twice. A large jagged red-orange box shows a person near a car at a road sign, caught between a phone that says a tornado warning ends at 4:15 and a TV crawl that says it ends at 4:00.
Same tool, two very different ways to disagree with itself

A new associate producer, two weeks into the job, messaged Rutendo directly: "quick one, the app says the warning ends at 4:15 and the crawl says 4:00, are those two different warnings?"

They weren't. Same county line, same National Weather Service alert, same input. One call had rounded the expiration time to the nearest quarter hour. The other had stated it exactly. Nothing was technically wrong. Nothing had ever been graded against the question of whether the two needed to say the same time.

Rutendo pulled up both channels at 4:02. The crawl still read "ends 4:00," already past its own stated cutoff, while the storm was still fifteen minutes from clearing the county line.

We didn't get the warning wrong. We got two right answers that disagreed with each other.

In the twenty minutes it took the assignment desk to sort out which time was right, the newsroom logged 46 calls asking the same question the new producer had just asked. Southridge read an on-air correction eight minutes later, stating the county's exact expiration time, word for word off the National Weather Service alert.

I want to say the problem is that one of the two channels was wrong. It wasn't. Both times were defensible readings of the same warning. The problem is that nothing had ever decided the two had to say the same thing, because nothing had ever been asked to.

Months earlier, when Cloudline first shipped, the team spent maybe ten minutes on the question of whether the warning card needed its own settings. Locking it down felt like extra engineering for a piece that had never once caused a problem. One shared setting, tuned for the blurb everyone thought about first, felt clean.

Rutendo pulled the warning card onto its own call the week after, locked to zero randomness, generated once and pushed identically to all four channels. Run the same April warning again, and the crawl and the app both say 4:15, because they're not two calls any more. They're one call, printed twice.

One design hands a viewer a fact she has to guess between. The other hands her the one fact that actually matters, said the same way twice.

And the thing I'd tell myself, if I could go back to the week we shipped Cloudline: we asked whether the writing was good. We never asked whether the same input, asked twice, owed anyone the same answer. On a calm Tuesday, it never mattered. On the one Thursday it did, we found out the hard way.

PICK, walked through on one warning card

This is a tradeoff about which piece of copy earns the deterministic pass, not a rule for every sentence Cloudline writes, so PICK is the tool.

P, position. Lock the storm-warning card to one deterministic call, at zero randomness, and reuse the exact same text across the TV crawl, the radio script, the app push, and the website banner. Leave the daily forecast blurb on its own separate, creative call.
I, impact. A wording mismatch on the daily blurb costs nobody anything, caught inside a normal shift if anyone notices at all, about 2 minutes. A mismatched warning reaches a live audience mid-storm: 46 calls in 20 minutes, an on-air correction, and a viewer near the county line left guessing which end time to trust.
C, cost asymmetry. A calm-day wording gap is loud the moment it happens and cheap to shrug off. A mismatched warning is invisible until the one live event where two channels happen to disagree, and by then someone has already acted on whichever screen they were looking at.
K, kill criteria. Drop the locked pass on the warning card the moment the identical text, repeated across a multi-day storm, starts driving people to stop opening it at all. Below that point, sameness is safety. Past it, sameness becomes its own kind of miss.
Knowledge spark: why not just have a producer proofread every warning against every channel, live? Because on the one afternoon it matters most, there isn't time. A real warning needs to publish in seconds, not sit while someone reads four screens side by side. The fix has to live in how the copy gets generated, not in a check added after.
Reader drop-off, if the exact same warning text repeats
Still above the kill line
Crosses it
Kill line: 40% open rate
0% 50% 100% 40% kill line 81% 77% 69% 58% 44% 31% 3rd repeat, worst real event so far 1 2 4 5 6
Every real storm Southridge has run through Cloudline has repeated the locked warning text at most three times before the event ended, and at three repeats the open rate is still a healthy 69%. The line only crosses the 40% kill line at a sixth repeat, a multi-day event longer than anything on record. Until a real event gets there, the locked pass stays exactly where it is.

Run PICK again, on a pharmacy counter

Anchorline Pharmacy runs the same question on refill alerts. Its AI drafts two kinds of message from the same fill record: a friendly "ready for pickup" text, and a "your refill window closes" text for controlled substances, sent by SMS and through the pharmacy's own app.

P. Lock the exact date and time in a controlled-substance refill deadline to one deterministic pass, reused identically across SMS and app. Leave the friendly "ready for pickup" reminder free to vary.
I. A patient who gets two slightly different "ready for pickup" texts loses nothing, maybe a shrug. A patient who reads two different refill-window dates for the same prescription might miss the real one, and for a controlled substance that can mean a lapsed prescription and a call back to the doctor.
C. The friendly reminder's mismatch is loud and cheap, caught by anyone glancing at two texts side by side. A wrong deadline is quiet until the day someone shows up to refill and finds out they were already a day late.
K. Drop the locked pass on any drug class whose refill window is wide enough that a same-day mismatch couldn't cause a missed dose, the same line Cloudline's team would use for a rule that no longer bites.

What I would leave alone, at the pharmacy counter The friendly "ready for pickup" text. It can read differently in the SMS and the app. Nobody's dose depends on how that sentence is worded.

Swap the trigger and it still runs

  • Speed: Cloudline answers in five seconds instead of twenty. Doesn't move the pick, because the position is about which piece of copy is allowed to disagree with itself, not how fast the draft arrives.
  • Cost: the locked warning-card call turns out to cost about the same in compute as just one of the four separate calls it replaced. Still doesn't flip it, the position was never really about compute.
  • The model gets better: if Cloudline's own repeat-call wording variance shrinks to nearly nothing on its own, fold the warning card back into the same setting the daily blurb uses. That's exactly the evidence that would let one setting do the whole job.

Where people run it wrong

  • Locking every piece of copy to zero randomness "to be safe," including the daily blurb, and ending up with a forecast that reads like a form letter nobody wants to open.
  • Treating a green pass/fail badge on the dashboard as proof the channels agree, without ever checking whether they used the same words for the same fact.
  • Fixing the mismatch channel by channel, so the crawl and the app each get their own patched-up rule instead of sharing one locked source.

If you are asked this cold

Say the reframe out loud before you answer yes or no. "Give me a second, I want to separate the pieces someone actually compares side by side from the ones nobody does before I say whether they all need to match." That's true, it's already stage three of the walkthrough, and it buys you the time to find the real asymmetry instead of guessing.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
Which framework fits this question, and what's the hardest step to nail?
Tap to flip
ANSWER
PICK, for a tradeoff. The hardest step is C, the cost asymmetry: naming why a mismatched warning card is hidden until a real storm hits, while a stale-sounding daily blurb is loud and gets caught in an afternoon.
2 · THE PERSON
Who owns Cloudline's consistency call, and what's she done for five years?
Tap to flip
ANSWER
Rutendo Chivayo, the product manager who has signed off on every release of Cloudline's warning card for five years, since before it wrote a word on its own.
3 · THE HABIT
What did Rutendo stop doing once Cloudline kept clearing every check?
Tap to flip
ANSWER
Reading all four channels side by side before anything went out. Then just skimming the crawl. Then just glancing at the pass/fail badge on the dashboard.
4 · THE ASYMMETRY
Name the two kinds of miss here and what each one costs.
Tap to flip
ANSWER
A calm-day wording gap on the daily blurb: unnoticed, costs about 2 minutes if anyone even tweaks it. A mismatched storm-warning card: caught by a new producer's question, 46 calls in 20 minutes, and an on-air correction eight minutes later.
5 · THE POSITION
State the pick in one sentence, the way you'd say it out loud.
Tap to flip
ANSWER
Lock the storm-warning card to one deterministic call, reused word for word across every channel, and leave the daily forecast blurb free to vary.
6 · THE NUMBER
The mismatched tornado warning drew ______ calls to the newsroom in the twenty minutes before Southridge issued an on-air correction.
Tap to flip
ANSWER
46. The crawl said the warning ended at 4:00, the app said 4:15, both drafted from the same input by two separate calls to the same model.
7 · THE KILL CRITERIA
What evidence would flip this position back the other way?
Tap to flip
ANSWER
Proof that the locked, identical warning text repeats so often in one multi-day storm that readers stop opening it. Southridge's own numbers put that around a sixth repeat, past anything they've hit so far.
8 · THE TRANSFER
Section 4 runs PICK again on a different product. Which one, and where does the position land there?
Tap to flip
ANSWER
Anchorline Pharmacy's refill texts. The locked pass sits on the exact date and time in a controlled-substance refill deadline, not on the friendly "ready for pickup" reminder.

Check yourself Score: 0 / 0

Multiple choice
1. Which of these is the actual mechanism behind this answer's pick?
  • A. Set every piece of copy Cloudline writes to zero randomness, including the daily forecast blurb.
  • B. Call the model once for the storm-warning card, at zero randomness, and push that same text to all four channels, while the daily blurb keeps its own separate creative call.
  • C. Have a producer manually compare all four channels by hand before anything airs, every time.
  • D. Cut Cloudline down to a single channel so there's nothing left to compare.
Show hint
Three of these either don't touch the one piece that actually broke, or throw away the time Cloudline was built to save.
Show answer
B. A locks down copy that never caused a problem, C undoes the whole point of the tool, and D throws out three working channels to avoid fixing one setting. Only B actually targets the miss that mattered.
Fill in the blank
2. Fill in the blank: before the fix, every piece of Cloudline's copy, warning card included, ran through the model at a shared sampling temperature of ______.
Show hint
The same setting that kept the daily blurb sounding fresh, applied to everything by default.
Show answer
0.7. Right for a piece that should sound a little different every day. Wrong for the one piece that got read on two screens at once during a live event.
True or false
3. True or false: this position requires every single piece of copy Cloudline writes, including the daily "sunny and 82" blurb, to come out identical every time.
  • True
  • False
Show hint
Think about who actually compares two copies of the daily blurb against each other.
Show answer
False. The position only locks the storm-warning card. The daily blurb keeps its own creative call on purpose, since a slightly different sentence about tomorrow's weather costs nobody anything.
Multiple choice
4. Why not just lock every piece of Cloudline's copy to zero randomness everywhere, to be safe?
  • A. It would take too much compute to run every channel that way.
  • B. Because the daily blurb isn't a piece anyone compares screen to screen, so locking it down trades a real everyday cost, a forecast that reads like a form letter, for a safety benefit that piece never needed.
  • C. Southridge's style guide forbids repeating any sentence twice in one week.
  • D. The model technically cannot run at zero randomness at all.
Show hint
Think about what sameness actually costs on a piece nobody is cross-checking.
Show answer
B. Determinism isn't free everywhere, it trades away useful variety on the piece where nobody was ever comparing two copies in the first place.
Short answer
5. If the mismatched warning had only drawn 5 complaint calls instead of 46, would the same position still hold? Walk through it.
Show hint
Think about whether the fix protects against the call count, or against the fact that identical inputs produced two different answers.
Show answer
Yes, still worth it. The 46 calls are evidence the mismatch happened and that people noticed, not the reason the fix matters. Even at 5 calls, the same input still produced two different facts about when it was safe to go outside. The call count tells you it's real. It doesn't set the bar for whether it's worth fixing.
Short answer, apply it yourself
6. Pick a product you use yourself. Name one piece of it where the same input should always give you back the exact same answer, and one piece where you'd actually rather it varied.
Show hint
Look for the one piece someone might screenshot and compare against a second screenshot of the same thing.
Show answer
Model answer: "A grocery delivery app: the estimated delivery window on my order confirmation and in my later text update should always match exactly, because I'll compare them if my order runs late. The 'here's what's fresh this week' banner on the home screen can say something different every time I open it, since nobody's cross-checking that against anything." Any answer works if you can name the piece someone actually compares against a second copy.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more