ConceptAdvancedShipping & Model Lifecycle / Model migration and version changes for users / #7
What contractual commitments should you avoid making about model behaviour?
The direct answer
Never write "will never," "always," or "100 percent" into a contract about what a model does. Write the commitment as a rate instead: a stated floor, measured on a named, audited eval set, checked on a fixed schedule, with retraining and service credits as the remedy when it slips, not automatic breach on one miss. A single miss inside that floor should trigger the recheck clause, not the termination clock.
Do this, in order
Replace absolute language with a threshold on a named eval set.Why: this one line decides whether the contract survives the model's normal, expected miss rate.
Name the eval set and the checking schedule in the contract itself, agreed jointly with the customer's own team.Why: a floor nobody can point to and re-run isn't a real floor, it's a promise dressed as one.
Write the remedy as retraining plus credits for a sustained miss, and save termination for a floor breached across more than one check.Why: it separates one expected miss from a real, ongoing failure.
Keep a human review step in the actual workflow, not just in the contract.Why: it catches the exact kind of miss an eval set won't catch on day one, a symptom message that reads like a routine follow-up.
Put the cost of a higher floor in writing before anyone asks for one.Why: pushing the floor from 99.0 to 99.8 percent roughly triples how many borderline messages get bumped to same-day triage, about 4 percent of all messages instead of 12. That's real nurse time, not a free upgrade.
Don't let a demo-stage reassurance set the real bar before any number is on paper.Why: an unwritten "it doesn't miss these" becomes the customer's baseline long before legal ever drafts a clause.
How to answer this, stage by stage
Seven moves. Say what a threshold actually protects before the story, or the answer sounds like a legal opinion instead of a decision.
1
Scope it to one company, one clause, one signature date
Say it like this
"Let's ground this. Oksana Farraday runs product for Vesper Health's scheduling assistant. Vesper is two days from signing a three-year contract with Cordillera Health, a forty-clinic network. Sales already sent over a draft clause: the assistant shall never fail to flag a red-flag symptom for same-day care. Oksana's the one who has to decide whether that sentence can stay in."
Why this works
A promise about model behaviour stays abstract until it's tied to one real clause, in one real contract, with a signature date already on the calendar.
2
Say your structure out loud
Say it like this
"I'll run this through LEAD. L is the real outcome the wording protects: a contract Vesper can actually keep for three years, not just at launch. E is the early signal, and you can read it off the page before anyone signs: is the clause a rate on a named eval set, or is it an absolute? A is how the wrong wording gets used against you. D is what even the right wording still can't promise."
Why this works
Naming the four letters up front tells the interviewer this is a contract-drafting decision with a real mechanism behind it, not a vague call for careful lawyers.
3
Reframe what the question is actually testing
Say it like this
"This sounds like a legal question. It's really a question about what a probabilistic system can honestly sign its name to. A person can promise to always check a badge at the door. A model can't promise to always catch the one message that reads like something else. The contract has to be written for the thing the model actually is, not the thing sales wishes it were."
Why this works
Moves the answer from "be careful with lawyers" to the real mechanism: a model has no zero setting, only a rate, and a contract that assumes zero is broken the day it's signed.
4
Give the decision straight
Say it like this
"Here's what I'd do. Kill the sentence that says 'shall never fail.' Replace it with a number: red-flag recall stays at or above 99.0 percent, measured every quarter against an eval set both sides sign off on. Two quarters below that floor triggers a mandatory retraining review and a credit. One miss inside the floor is not, on its own, a breach, because 99.0 percent already means it will miss sometimes, by design."
Why this works
Names the exact number, the exact eval mechanism, and the exact line between an expected miss and a real failure, instead of a general call for more testing.
5
Prove it with the number that shows why "never" breaks
Say it like this
"Here's why that matters. Vesper's own red-flag recall was 97.1 percent at launch, climbed to 99.3 percent after a retrain, dipped to 98.0 percent the quarter they migrated the underlying model, then settled at 99.0 percent by the time Cordillera's pilot started. It has never once touched 100, and over a three-year contract with two more model migrations already planned, it isn't going to. A clause that says 'never' was broken the day it was signed, whether anyone had noticed yet or not."
Why this works
One real trend line, including a migration dip, does more work than a paragraph arguing that models are imperfect in general.
6
Name the abuse path
Say it like this
"Here's how the wrong wording gets used. Four of Vesper's other enterprise accounts signed contracts with 'will never miss' language. Three of those four filed a formal breach dispute the first time the assistant missed one case, even though recall never left the 99-percent range. Zero of the five accounts on threshold wording did that. Same product, same miss rate, completely different legal exposure, based only on which sentence was in the contract."
Why this works
Shows the exact mechanism by which vague or absolute wording gets exploited, a real split across real accounts, not a hypothetical.
7
Close on the hard limit, and say what you'd leave alone
Say it like this
"And here's the limit, worth saying to Cordillera directly. Even at 99.0 percent, this system will miss sometimes. That's what the number means. What I can promise is that we catch it fast, we retrain when the rate slips, and we keep a nurse reviewing new bookings so a miss doesn't reach a patient before someone looks twice. What I wouldn't touch: our uptime commitment and our data-handling language can stay absolute. No patient data leaves this system, ever, encrypted at rest, always, because that's something our code either does or doesn't do. There's no probability distribution over whether a file got encrypted."
Why this works
Shows judgment instead of blanket caution about "AI risk." An absolute promise is fine exactly where the system is deterministic, and wrong exactly where it isn't.
Let's learn
What happens the first time a written promise about a model turns out to be a promise nobody could actually keep?
Say a hospital network lets patients book their own appointments by texting an assistant. It reads what a patient types, works out which specialist they need, checks whether that provider is in their insurance network, and if anything in the message sounds urgent, it's supposed to bump them to same-day care instead of the next open slot.
Before anything like this existed, patients called the front desk, and a receptionist decided, by ear, whether a message sounded urgent enough to interrupt the schedule. Now the assistant does that first pass, in under two minutes, for almost every message that comes in.
It doesn't catch every red flag. No version of it ever has. On Vesper's own audited eval set, 2,600 clinician-written symptom scripts, red-flag recall started at 97.1 percent and has climbed since, but it has never touched 100.
Knowledge spark: what's an eval set?
A fixed stack of test messages, written and graded ahead of time by real clinicians, used to check the model the same way every quarter. Run it again next quarter, get a number you can actually compare to this quarter's.
Here's the turn. The extra misses aren't really the problem, because at 99.0 percent there are always going to be some. The problem is what Cordillera's legal team reached for the moment they saw one: not a number, a sentence. Their draft said the assistant "shall never fail to flag" a red-flag symptom, with a termination-for-cause clause sitting behind it, live from the day both sides signed.
Vesper's red-flag recall, by quarter, across two model changes
Red-flag recall, audited eval setContracted floor and current quarter
The line never touches 100, not even after a retrain, and it visibly dips the quarter the base model changes. A clause built on "never" was already false before Cordillera ever saw a draft.
The near miss did not break the contract. The sentence promising there would never be one did.
The near miss itself was small. A patient described chest tightness and shortness of breath after climbing the stairs at home. Because the same message also referenced an existing cardiology follow-up, the assistant read it as routine and offered a slot two weeks out. A triage nurse doing her afternoon check of new bookings, a habit Cordillera had kept even after the assistant went live, caught it that same afternoon and called the patient in. Nobody was hurt.
What almost got hurt was the contract. Once Cordillera's legal team saw the incident log, they wanted a sentence with no room in it at all, and a sentence like that turns every future in-band miss, the kind the assistant was always going to have, into grounds to walk away.
Accounts that filed a breach dispute after one miss, by contract wording, same underlying recall rate
3 of 4
"Will never miss" wording
0 of 5
Threshold, eval-set wording
Same product, same recall rate in the 99-percent range the whole time. The dispute wasn't about the miss. It was about which sentence was sitting behind it.
The clause was telling the truth about what it said. The waiting room was telling the truth about what actually happened. Both were true the same afternoon.
The choice I'd take back
Months earlier, on the first sales call, Cordillera's ops lead asked Chidi Vantwisk how often the assistant misses a red flag. He told her, honestly meaning it as reassurance, "it just doesn't miss these." Nobody wrote a number down that day. That unwritten line sat there, unquestioned, until legal drafted the clause it had always implied. I'd take that back: the first time red-flag accuracy came up on a call, I'd have put a number in the room myself, not left a reassurance to harden into a promise on its own.
What I'd leave alone. Vesper also promises 99.9 percent uptime and that no patient message ever leaves the system unencrypted, in either direction. Both stay absolute. Uptime is already a rate, and the encryption line describes something the code either does or doesn't do every single time, with no model judgment involved. Only the model-behaviour clauses needed rewriting.
The lesson. A sentence that promises zero mistakes isn't a stronger commitment than a rate. It's a weaker one, because it's already false the day it's signed, and it hands the other side a trigger for something the system was always going to do eventually. A rate, checked on a schedule, is the only version of the promise that survives contact with a real quarter.
Now here is the same thing as a story
Read this version when you've got the extra four minutes. The short one above is what you'd say in an interview. This is why it's true.
Oksana Farraday had closed six enterprise healthcare contracts before this one, and she had a rule she'd learned the hard way: read the redline before the call where everyone celebrates it. She was two days from Cordillera Health signing, forty clinics, three years, the biggest account Vesper had ever chased.
The pilot had gone well. Six weeks, real patients, real bookings, red-flag recall holding at 99.0 percent. The first two weeks, Cordillera's triage nurses reviewed every new self-scheduled booking by hand every afternoon, the way they'd reviewed the front desk's work for years. By week four, with nothing wrong yet, they'd cut that down to spot-checking about one in five.
The chest-tightness message happened to land in the one-fifth they still checked. That was luck, not design, and Oksana knew it the moment she heard about it.
Two days before signature, Cordillera's general counsel sent over a redline. One new sentence, buried in the service-level section: the assistant shall never fail to flag a red-flag symptom for same-day care. Behind it, a termination-for-cause clause, live from day one.
Oksana pulled the call transcript from the first sales conversation, four months earlier. Cordillera's ops lead had asked Chidi, plainly, how often the assistant gets it wrong on something urgent. Chidi's answer, word for word: "Honestly, it just doesn't miss these." Nobody in that room had a number. Nobody asked for one. The reassurance just sat there, and four months later it had grown into a sentence with a lawyer's signature line under it.
Nobody had lied on that call. It just turned out that "it doesn't miss these" and "it will never miss one" are not the same sentence, and only one of them can survive three years.
Oksana considered, for about an hour, just asking Chidi to soften the sales deck and letting the redline stand, because the recall rate really was excellent and the odds of a dispute in any given quarter were low. She rejected that. A low odds of a bad quarter isn't a zero, and a contract written for the good quarters only breaks in exactly the quarter it was supposed to protect against. She also considered proposing a blanket liability waiver instead, no accountability for any AI-related miss at all. She rejected that too, faster: Cordillera's counsel would never sign a three-year deal with no real commitment behind it, and asking would likely kill the negotiation outright.
What she actually did was call Cordillera's counsel directly. "Before we talk about never," she said, "I want to show you a number. Our recall on red flags has been at or above 99.0 percent for two straight quarters, including through a full model migration. I can commit to that number, in writing, checked every quarter against an eval set your own clinical team helps write. I cannot commit to zero, because nobody honestly can, and a sentence that says zero is worse for you than a number, because it gives you nothing real to hold us to once it's already broken."
The redline changed. 99.0 percent floor. Quarterly check against a jointly maintained eval set. Two consecutive quarters below floor triggers a mandatory retraining review and a service credit. One miss inside the floor, on its own, changes nothing about the contract.
Eleven months in, it got tested. Vesper migrated to a newer base model, and for one quarter red-flag recall dipped to 98.6 percent. Under the old wording, that dip would have sat behind a live breach clause the day it happened. Under the rewritten one, it triggered a five-day retraining review and a credit on Cordillera's next invoice. The contract was still standing at month eighteen, still is now, and Cordillera's clinical team still helps write the eval set every year.
One sentence handed Cordillera a trigger for the day the model did something every model eventually does. The other handed both sides a number they could actually watch.
And the thing I'd tell my past self, back on that first sales call when "it just doesn't miss these" first got said out loud: the moment somebody says a model is perfect, write down the number that proves it isn't, before someone else writes down the sentence that assumes it is.
LEAD, so the contract survives the miss it's going to have
This sounds like a legal-drafting question. Underneath, it's still asking what a probabilistic system can honestly sign its name to. That's LEAD, run on a contract clause instead of a dashboard.
L, link. The real outcome a well-written commitment protects. Not a sales deck that says the assistant is flawless. A contract Vesper can actually keep for its full term, as the model gets retrained and migrated, not one that's broken the moment it behaves probabilistically even slightly outside what a sentence like "never" assumed. → Here, that's whether the three-year Cordillera deal survives Vesper's own model roadmap, two more migrations already planned, instead of breaking on the first expected miss.
E, early signal. Whether a proposed commitment is phrased as a threshold and a named eval set, or phrased as an absolute guarantee. You can read this straight off the term sheet, weeks before any dispute happens. → "99.0 percent, checked quarterly, on eval set X" survives a bad quarter. "Shall never fail" doesn't, and you can tell which is which before either side has seen one.
A, abuse. How vague or absolute contract language gets exploited. A customer cites one in-band miss as a contract breach, when the real commitment should have been a rate, not a promise of zero mistakes. → Three of Vesper's four "will never miss" accounts filed a breach dispute after exactly one miss. Zero of five threshold accounts did, for the same underlying performance.
D, decision. What a well-written, threshold-based commitment genuinely cannot promise: perfect behaviour on every single request. Only a bounded, monitored failure rate. → Oksana can promise Cordillera 99.0 percent, checked every quarter, and a fast fix when it slips. She still can't promise the next message is the one it catches.
The check that proves a clause is real
Ask, before you sign anything: "if this promise gets tested by one bad output next quarter, does the contract survive?" If the honest answer is no, the sentence isn't a commitment. It's a guess wearing a commitment's clothes.
And if you want to be sure it really works, try it somewhere else
Kestrelmoor Insurance runs an AI assistant that triages incoming auto-claims, recommending which ones get fast-tracked for payout and which get routed to a human adjuster. Nadja Sundstrom, who leads that product, is mid-negotiation on a five-year distribution deal with Ashford Mutual, a reinsurer that wants language guaranteeing the assistant "shall never deny a claim that should have been approved."
L. Whether Kestrelmoor can actually stand behind the deal for its full five-year term, across at least one planned fraud-model retrain a year, instead of it breaking on the first wrongful denial anyone spots.
E. Readable in the draft clause itself: a false-denial rate with a named holdout set and a monthly check, or the word "never."
A. Ashford's earlier distribution deal with a different vendor carried "never" wording. One wrongful denial, out of a very large claim volume, let Ashford terminate for cause, even though that vendor's real false-denial rate over the year was 0.4 percent, industry-leading. Kestrelmoor's own draft was heading the same way before Nadja caught it.
D. Even a strong threshold, say a 0.5 percent false-denial rate checked monthly, doesn't promise any one policyholder's claim gets it right. It promises the rate holds, and it gives Ashford a real number to hold Kestrelmoor to, instead of a sentence that breaks on the first unlucky claim.
Knowledge spark: what's a false denial here?
A real, valid claim the assistant recommends for denial or extra review it didn't need. The claimant doesn't find out why. They just get a slower, harder process for a claim that should have gone straight through.
Nadja rewrote Ashford's draft the same way Oksana rewrote Cordillera's: a 0.5 percent false-denial floor, a jointly maintained holdout set of adjudicated claims, monthly checks, retraining and credits below floor, termination reserved for a floor breached across three straight months. Ashford signed it in a week, faster than the original "never" draft had been moving.
Same product, same single miss, two different documents. Only one of them was still standing the week after it happened.
Swap the trigger and it still runs
The model gets more accurate every quarter. Doesn't help. A model climbing toward 99.5 percent is still not at 100, and an absolute clause is still broken the same way.
The contract only runs one year instead of three. Doesn't help either. One model migration inside even a single year is enough to test a "never" clause.
The customer never actually has a single miss during the whole term. Still needs the threshold wording. Nobody signing a multi-year deal gets to assume the lucky run continues.
Where people run it wrong
Letting a sales call's reassurance sit unwritten until legal drafts something worse from it.
Writing a floor with no named eval set, so nobody can agree later on what actually got measured.
Treating every AI-related clause the same way, when the parts of the system that aren't the model, uptime, encryption, can carry a firm promise just fine.
How to use it live
Say the split out loud, early: "Before we talk numbers, I want to separate two things: what this system does on a good day, and what it's contracted to do on its worst normal day. Those are different sentences." That's not stalling. It names the actual gap between a demo and a document, and it buys you room to build the real threshold instead of signing whatever's already been said out loud on a call.
Flashcards (click a card to flip it)
1 · THE FRAMEWORK
Which framework fits this question, and what does each letter stand for here?
Tap to flip
ANSWER
LEAD. L is the real outcome, a contract Vesper can keep for its full term as the model changes. E is the early signal, whether the clause is a rate on a named eval set or an absolute. A is how it gets gamed, a customer citing one in-band miss as a breach. D is what it can't promise: perfect behaviour on every request, only a bounded, monitored rate.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Oksana Farraday, who runs product for Vesper Health's scheduling assistant, and rewrote the red-flag clause in Cordillera Health's contract two days before signature.
3 · THE HABIT
What let the near miss almost slip through unnoticed?
Tap to flip
ANSWER
Cordillera's triage nurses had cut back from checking every new booking to spot-checking about one in five, because the pilot had gone clean for weeks. The chest-tightness message happened to land in the checked fifth, by luck.
4 · THE EARLY SIGNAL
What's the E step here, in one line?
Tap to flip
ANSWER
Whether the clause is written as a rate on a named, audited eval set, checked on a fixed schedule, or as an absolute like "never" or "always." Readable in the term sheet, before any miss ever happens.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Letting Chidi's demo-call reassurance, "it just doesn't miss these," go unwritten. It sat unquestioned for four months and hardened into the literal contract sentence Oksana had to unwind two days before signature.
6 · THE NUMBER
The rewritten contract set a red-flag recall floor of ______ percent, checked ______. Three of Vesper's four "never miss" accounts filed a dispute after one miss; ______ of five threshold accounts did.
Tap to flip
ANSWER
99.0 percent, checked quarterly. Zero of five threshold accounts filed a dispute, for the same underlying recall rate as the accounts that did.
7 · THE REPLAY
Eleven months in, recall dips after a model migration. What changes with the rewritten clause?
Tap to flip
ANSWER
Recall dips to 98.6 percent for one quarter. Instead of a live breach clause, it triggers a five-day retraining review and a service credit. The contract is still standing at month eighteen.
8 · THE TRANSFER
Section four runs this same question again for a different product. Which one, and what does the same fix look like?
Tap to flip
ANSWER
Kestrelmoor Insurance's auto-claims triage assistant. Nadja Sundstrom rewrote Ashford Mutual's "shall never deny" clause into a 0.5 percent false-denial floor, checked monthly against a jointly maintained holdout set.
Check yourself Score: 0 / 0
Fill in the blank
1. Oksana rewrote Cordillera's "shall never fail" clause into a red-flag recall floor of ______ percent, checked ______, against a jointly maintained eval set.
Show hint
This is the exact number from the direct answer and the walkthrough's stage 4.
Show answer
99.0 percent, checked quarterly. Two consecutive quarters below that floor trigger a mandatory retraining review and a credit, not automatic breach on a single miss.
Multiple choice
2. Why can't Vesper honestly promise the assistant will never miss a red-flag message?
A. Because red-flag messages are rare and hard to collect enough data on.
B. Because the model is probabilistic, so it always keeps some non-zero failure rate, and that rate can shift when it's retrained or migrated.
C. Because Cordillera's clinics see too much volume to check every message.
D. Because Vesper hasn't finished testing the assistant yet.
Show hint
Think about what "probabilistic" actually means for a sentence written into a contract.
Show answer
B. The others describe operational or timing issues, not the real reason. A probabilistic system has no zero setting, only a rate, and that rate is exactly what moved when the base model migrated in Q3.
True or false
3. True or false: because Vesper's red-flag recall stayed at or above 99.0 percent all year, none of its enterprise accounts ever had a dispute over a miss.
True
False
Show hint
Check which accounts had disputes, and what kind of contract wording they'd signed.
Show answer
False. Three of four accounts written with absolute "will never miss" language filed a formal dispute after exactly one miss, even though the underlying recall rate never left the 99-percent range. The dispute was about the sentence, not the number.
Multiple choice
4. Which of these Vesper commitments is fine to leave in absolute, deterministic language, even inside an AI product's contract?
A. "The assistant will always correctly identify a red-flag symptom."
B. "The assistant will never route a patient to an out-of-network provider by mistake."
C. "No patient message ever leaves the system unencrypted."
D. "The assistant's specialist recommendation will always match what a clinician would choose."
Show hint
Look for the one commitment that isn't a judgment call the model makes.
Show answer
C. Encryption is something the code either does or doesn't do every time, with no model judgment involved. The other three describe probabilistic model behaviour, which belongs in a rate, not an absolute.
Short answer, apply it yourself
5. Think of a product you use that comes with some kind of promise about what it will always or never do. Where would you rewrite that promise as a rate instead, and what would the eval set be?
Show hint
Look for a promise that assumes zero mistakes instead of stating a checked number.
Show answer
Model answer: A ride-hailing app might promise a driver "will always arrive within the quoted window." In reality that's a rate: most rides land within a few minutes of the estimate, checked against actual GPS logs each month, not a guarantee every single ride hits it. Writing it as a rate is honest, and it's still something a rider can hold the company to.
Fill in the blank
6. At month 11, after Vesper migrated its base model, red-flag recall dipped to ______ percent for one quarter. Under the rewritten contract, that triggered a ______-day retraining review and a service credit, not a breach notice.
Show hint
This is the number from the end of the story section, not the launch-quarter numbers.
Show answer
98.6 percent. A five-day retraining review. Under the old "shall never fail" wording, that same dip would have sat behind a live breach clause the day it happened.
If the interviewer pushes back
Why this works
Tests whether you'll turn a real near miss into a contract fix instead of a legal argument, and whether you can tell a rate from a promise the moment you read it, before a dispute ever happens.
Follow-up traps
"What if 99.0 percent still isn't good enough for something this serious?"Raise the floor, don't remove the rate, and say what it costs: a 99.8 percent floor roughly triples how many borderline messages get bumped to same-day triage, about 12 percent of all messages instead of 4. That's a real trade Cordillera can choose, not a free upgrade.
"Isn't a threshold just a way to dodge accountability for mistakes?"No, the opposite. An unenforceable "never" clause gives a customer nothing to hold you to, because it's already false the day it's signed. A named eval set and a fixed schedule is the part they can actually audit.
If pressed
The eval set itself needs versioning too. Once Cordillera's clinical team adds new red-flag scripts drawn from real near misses, like the chest-tightness case, the recall number on the old eval set and the new one aren't directly comparable. The contract also covers how the eval set gets updated without resetting the floor's history to zero.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.