InterviewAdvancedModel Fluency & the AI PM Role / What changes when the product is probabilistic / #21
Walk me through how you would convince a skeptical enterprise buyer that your AI feature is reliable enough to trust.
PICK · a live vendor-risk review of a contract-reading AI, at a steel manufacturer
Indemnis is Wrensmoor's AI tool for reading enterprise supply and vendor contracts: it flags the clauses that could actually cost real money, before anyone signs. Zell Kase owns its trust and reliability story. Ninety days into a paid pilot, Kestrevale Steel's head of vendor risk, Cateline Feathering, closes her laptop and asks the one question that decides whether the deal happens: why should she let a model anywhere near a contract that could cost her company millions.
The direct answer
Don't tell Cateline Indemnis is accurate. Tell her it's measured, and tell her where it's still allowed to be wrong. On 240 real disputed contracts, Indemnis misses a genuine liability or indemnification issue about 2.5 percent of the time, and no contract worth more than $2 million, and no liability or indemnification clause at any price, ever clears without a lawyer's own eyes on it, no matter what Indemnis's own score says. Then say plainly what number would make you pull it back yourself, before she has to ask.
Do this, in order
Open with the position, not a reassurance: measured and gated, never claimed perfect.Why: "trust us" is a slogan; "here's the number and here's the gate" is a claim someone can check.
Name who actually eats each kind of wrong, in real terms, both directions.Why: an unnamed cost stays abstract, and an abstract cost is easy for a skeptical buyer to wave away or wave through.
Gate every high-value contract and every liability-shaped clause on a person, no matter the score.Why: the miss that can't be undone is the one worth designing against; a missed lunch-order form isn't.
Show the eval set and the measured rate before Cateline asks for them.Why: a number you volunteer reads as confidence; a number you get dragged into giving reads as hiding.
Say what would make you pull Indemnis back, unprompted.Why: only a claim that can be proven wrong is a real claim; a claim with no way to fail is a pitch, not a case.
Pin the model version, and promise a re-check before any swap.Why: every number above is only true of one exact model; a silent upgrade breaks the whole case without anyone deciding to.
How to answer this, stage by stage
Nobody is grading whether you can sound confident. They're grading whether you can hold up under a reviewer whose whole job is to find the hole in your story, and whether you'd hand her the hole yourself before she finds it.
1
Put a real reviewer in the room, not a hypothetical one
Say it like this
"I'll answer this the way I'd actually run it, live. Indemnis is Wrensmoor's tool for reading enterprise contracts and flagging what could cost real money. Say I'm ninety days into a paid pilot with a steel manufacturer, and their head of vendor risk has just asked me straight: why should she trust a model near a contract worth millions."
Why this works
A named reviewer with a real question keeps the whole answer honest; you can't hand-wave at a specific person the way you can at "stakeholders."
2
Announce the shape before making a single claim
Say it like this
"Before I answer, here's how I'm going to answer it. My actual position, first, no softening. Then who really pays if I'm wrong, on both sides. Then which kind of wrong should scare her more, and why. Then exactly what would make me tell her not to trust it yet."
Why this works
Two seconds of structure tells a skeptical reviewer she's watching a method, not being talked at until she gives up asking questions.
3
Say what "reliable" actually has to mean here
Say it like this
"Reliable doesn't mean Indemnis never gets a clause wrong. Nothing that reads a contract at scale gets every clause right, and anyone who tells you otherwise hasn't been asked a hard question about it yet. Reliable means we know exactly how often it's wrong, we know which kind of wrong actually hurts you, and we've built the system so that kind of wrong never reaches you unwatched."
Why this works
This reframe is the whole answer in miniature. Skip it and everything after sounds like a longer version of "trust us."
4
Give the position, committed, before any evidence
Say it like this
"Here's my actual position. Trust Indemnis to clear the contracts where a miss costs you an afternoon. Never trust it alone on the ones where a miss costs you millions, and you don't have to, because those never clear without a lawyer's own eyes, automatically, no matter what score Indemnis gives them."
Why this works
This is the direct answer, said out loud, before Cateline has to dig for it through ten minutes of hedging.
5
Name who actually feels each kind of wrong
Say it like this
"Two people carry this, and they carry different things. Your contracts manager reads Indemnis's clear verdict and moves on. If it's wrong once on something real, she stops trusting it on everything, even the contracts it gets right, and goes back to reading all eleven hundred of them by hand. And you, if you get spooked and force full manual review across the board just in case, you never get the thing you're actually paying for, and your legal team is right back where they started."
Why this works
Naming both people keeps this from turning into a one-sided sales pitch about how great the model is.
6
Give the asymmetry, with the real numbers behind it
Say it like this
"Here's the number, and here's why it's not the whole story. Measured against 240 real disputed contracts, Indemnis misses a genuine liability or indemnification issue about 2.5 percent of the time. That number should scare you a little on a forty thousand dollar order form. It should not reach you at all on a two million dollar supply deal, and it doesn't, because value and clause type route straight to a person, every time, whatever Indemnis's own confidence says."
Why this works
This is the hardest move in PICK. Naming which kind of wrong is cheap and which is expensive, in the reviewer's own terms, is what turns a rate into a real case.
7
Prove it with the near miss, compressed to four sentences
Say it like this
"And I'll tell you the one that almost got through, because you'd find it eventually anyway. Ninety days ago, a supply contract scored 96 percent sure overall, but one clause buried in an appendix, alone, would have scored 61. It capped our damages at replacement value only, on a deal where a real defect could run past four million. A contracts manager caught it three days before signing, on a routine check, not luck, and that's the exact gap the per-clause gate exists to close now."
Why this works
A real near miss, told plainly and without spin, makes the case sound like risk management instead of a demo reel.
8
Say the kill line unprompted, then close
Say it like this
"Last thing, and I'll say it before you ask. If our next batch of confirmed real cases shows Indemnis missing liability issues more than twice as often as it does today, past 5 percent, I'll tell you myself to stop trusting it there until we fix it. If we ever change the model underneath it, you get a fresh test result before the new one goes live on your account, not after. So: measured, gated on the wrong that can't be undone, and I've already told you what would make me take it back."
Why this works
A pitch that names its own failure condition, unprompted, is the difference between a case and a sales pitch, which is exactly what this question is testing.
Let's learn
Indemnis reads a contract before anyone signs it and tells a legal team, in plain words, which clauses could actually hurt them.
One contract reads like one document. It's really five different kinds of promise stacked together, and only some of them can bankrupt you if they're wrong.
Before Indemnis, Kestrevale Steel's six in-house contract attorneys and paralegals read every supply and client contract by hand, about 1,100 of them a year, at roughly 3.5 hours each. That's real cost: about $330 a contract once you count the attorney time, and a backlog that made vendor onboarding take close to two weeks. With Indemnis reading first and flagging what needs a real look, the same team clears most contracts in under an hour, and the average cost per contract drops to about $95, blended across the ones that clear fast and the ones that still get a full attorney review.
Five steps. The fourth one, checking value and clause type separately from the score, is the step the whole trust case actually turns on.
Knowledge spark: what's a false negative here?
Indemnis reads a clause and calls it fine. It wasn't. It missed a real problem. That's the mistake that matters most for a contract tool, because the clause still goes out the door looking clear.
Measured against 240 real contracts that were later confirmed as disputed, in court or arbitration, Indemnis misses a genuine liability or indemnification issue about 2.5 percent of the time. For comparison: one attorney doing a single fast first read alone, no second check, misses about 4.1 percent of the same set. Indemnis is already a little better than a rushed human on its own. That's not why Kestrevale should trust it.
Three separate tripwires, any one of them enough on its own. A contract only clears itself when none of the three fire.
Here's the turn. The scary version of this was never Indemnis missing a clause now and then. It's what happens after: if Cateline's team gets spooked by one story like that and forces full manual review across every contract regardless of size, Kestrevale is right back to $330 a contract and a two week backlog, and the entire reason to buy Indemnis disappears.
We were never scared of Indemnis missing a clause. We were scared of Cateline turning it off.
Cost per contract, full manual review vs Indemnis with the gate
Full manual review, per contractIndemnis with the gate, per contract
Across 1,100 contracts a year, that gap is worth about $258,000. Panicking and turning the gate into full manual review for everyone spends that whole number back, on top of the two week backlog it took months to remove.
The choice I would take back
When Wrensmoor and Kestrevale scoped the pilot's rules, the plan was one blended confidence score for the whole document, gate anything under 90 percent, ship the rest. That made sense for a two week setup with a launch deadline. It stopped making sense the moment a document could sit at 96 percent overall while one clause inside it, alone, was a coin flip.
What I would leave alone: routine purchase orders under $50,000 written on Kestrevale's own standard template don't need any of this. Indemnis clears those fast, and it should, because the boilerplate there hasn't changed in years and nobody has ever disputed a line in it.
The lesson: a reliability story that only reassures isn't evidence. A buyer who's heard a pitch before can tell the difference between "trust us" and "here's the number, here's the gate, and here's what would make me take it back." Only one of those survives a hard question.
Now here is the same thing as a story
The short version above is what you actually say in the room. Read this one when you want to feel why a clause on page fourteen of an appendix nearly mattered more than the whole rest of the deal.
Osaze Marnier has managed Kestrevale Steel's supply contracts for nine years, and she has one habit nobody ever had to teach her: she always flips to the last schedule in any appendix first, because that's where the surprises hide. Payment terms and delivery windows sit up front, easy to write, easy to check. The strange language always ends up on page twelve or fourteen, tucked in after everyone's attention has already left the room.
Indemnis arrived for the pilot in the spring, and for the first two months it was the best thing that had happened to her team's mornings. It read a contract overnight and handed back a scored draft by eight, flagging the handful that needed real attention out of the dozens that came in each week. Osaze kept her old habit at first, reading every appendix in full anyway, the same way she always had, checking Indemnis's high scores against her own eyes.
The scores kept agreeing with her. Week after week, nothing she found contradicted what Indemnis had already flagged clean. So the full read thinned, a little at a time, the way habits do when nothing punishes them. By week eight, she was only reading appendices in full on contracts Indemnis scored below 90 percent. Above that, she skimmed the summary and moved on. Nobody decided this on purpose. It just kept being fine, until the week it wasn't.
Contract 187 was a supply agreement addendum with Kestrevale's largest single alloy vendor, worth just under $1.9 million, a hair under the value gate that didn't exist yet. Indemnis scored the whole document 96 percent sure. Osaze almost let the appendix go with a skim, the way she'd been doing for six weeks by then. She caught herself flipping to section 14.3 anyway, out of the decade-old habit her system had quietly stopped asking of her.
A ten year old habit, not luck, closed a gap the scoring system hadn't been built to see yet.
Section 14.3 read, in the flat language addenda always use, that Kestrevale's remedy for a defective shipment was limited to replacement value only, notwithstanding any other provision in the agreement. On a deal that size, a genuine defect claim downstream, say a batch of alloy that failed in a customer's own product, could easily have run past $4 million in consequential damages Kestrevale had never meant to accept. Scored on its own, that one clause would have come back at 61 percent. Blended into the rest of a clean, well-drafted document, it disappeared into a 96.
Ninety six percent sure was true. It just wasn't true about the one clause that mattered.
Osaze caught it three days before signature, called the vendor's counsel, and got the clause rewritten to match Kestrevale's standard cap. Nobody outside her own team ever knew it happened. She wasn't lucky. She was doing the thing she'd always done, on a system that had quietly stopped requiring it of her.
The decision Zell would take back traces to a kickoff call ten weeks earlier. Someone on the scoping call had asked whether the gate should route by clause type as well as by score. The answer, reasonable at the time, with a launch date two weeks out, was to start with one blended number and add more gates later if they turned out to matter. That was a fine call for a two week setup. It was never built to survive a document that could be genuinely excellent everywhere except one paragraph.
Run Contract 187 again with the gate rebuilt the way it works now, scoring liability, indemnification, and consequential-damages language on its own, separate from the rest of the document, regardless of the overall score. Section 14.3 comes back at 61 percent, well under the 92 percent line, and routes to an attorney automatically, in the same overnight run that produced the draft. Nobody needs to remember to flip to page fourteen. The system does it before the document ever reaches a desk.
That's the story sitting behind the ninety day mark. And the ninety day mark is where Cateline Feathering, Kestrevale's head of vendor risk, closes her laptop and asks Zell the question this whole page is built to answer. Zell doesn't open with the accuracy number. He opens with the position, tells her about Contract 187 without being asked, and ends on the exact number that would make him agree with her if she decided not to trust it. Cateline asks two follow-up questions. She gets straight numbers back for both. Kestrevale signs the three year deal two weeks later, with the gate written into the contract as a requirement, not a feature.
The whole answer to this question, in one picture. A confident average is not the same thing as a confident answer about the one part that could actually hurt you.
What Zell would tell himself, back on that kickoff call: starting with one blended score wasn't careless. It was the sensible call for a two week setup with a deadline. It just wasn't a call anyone had agreed to revisit once the model got good enough that nobody had a reason left to check behind it.
PICK: the four things Cateline needed before she'd believe any of it
Not a way to sound confident under pressure. PICK is what forces a real commitment about which number this claim actually stands on, and which kind of wrong Cateline should actually be afraid of.
PPosition. The real claim, in one sentence, before any reasoning.
We don't promise Indemnis is always right. We promise it's measured against real disputed contracts, monitored clause by clause, and it can never clear a liability or indemnification section on its own once the stakes are high enough to matter.
Say the position first, addressed straight to the person across the table. An answer that buries its claim inside the reasoning has already lost a skeptical reviewer.
IImpact. Who feels each kind of wrong, in the reviewer's own terms.
If Indemnis misses something real and it reaches a contracts manager unwatched, she stops trusting it on everything, not just the contract it got wrong, and Kestrevale loses every hour it was ever saving. If Cateline panics and mandates full manual review of every contract regardless of size, Kestrevale's legal team goes right back to $330 a contract and a two week backlog, and the entire reason to buy Indemnis disappears.
Naming both people, not just the model's own error rate, is what keeps this from turning into a one-sided pitch about how good the tool is.
CCost asymmetry. The heart of it.
A missed clause on a routine order under $50,000 is the cheap mistake: it costs an afternoon and a reread, absorbed inside the team's normal week. A missed liability or indemnification clause on a multi-million dollar deal is the expensive one, and it's the one Cateline should actually fear, because it can turn into a claim nobody can walk back once material has already shipped. That's exactly why value and clause type route to a person automatically, whatever Indemnis's own confidence says: the expensive mistake never gets the chance to happen unwatched.
This is the step that earns the pick. Anyone can say a tool is reliable. Naming which failure is the one worth designing against is what makes the case survive a follow-up question.
One box is small on purpose. The other one is drawn the size it actually costs.
KKill criteria. What evidence flips the pick.
If the next quarterly batch of confirmed real disputes shows Indemnis missing liability or indemnification issues past 5 percent, roughly double today's rate, that's a real signal, not noise, and the gated review queue gets widened until it's fixed. If the underlying model version ever changes without a fresh 240-contract test run and a shared report before it goes live on Kestrevale's account, that alone is grounds to distrust the new version, regardless of what its own number says, because at that point nobody actually knows the real rate yet.
A pick with no way to be proven wrong is a hope, not a claim. Naming the exact bar, before Cateline has to ask for one, is what makes this a real case instead of a promise.
The kill line, charted: false negative rate on liability and indemnification clauses, by quarter
Flat and well under the line for four straight quarters. The line only matters because Zell said the number out loud before Cateline asked what would move it.
Three things worth stating directly, since this is where the real judgment sits. The alternative Zell's team considered, and rejected, was simply raising the overall blended confidence bar higher, say from 90 to 98 percent, instead of building separate per-clause gates. It lost, because Contract 187 scored 96 percent blended. A higher blended bar would still have let it through; the flaw was never the height of one number, it was scoring the whole document as a single average when one paragraph inside it needed its own answer. The AI-specific failure worth naming by name is exactly that: a document-level confidence score can look excellent while one specific clause inside it is genuinely a coin flip, and nothing about a blended average tells you which clause that is. The guardrail is a risk-class rule, not a smarter threshold: liability, indemnification, and consequential-damages language get scored and gated on their own, at any value, and any contract over $2 million gets a person's eyes regardless of what every clause inside it scored. And the trade-off is real and accepted on purpose: a gated contract takes two to three business days for an attorney's own read instead of clearing overnight, which is slower and costs real attorney hours. Kestrevale accepts that cost on purpose for the slice of contracts where a miss can't be undone once a shipment has already gone out the door.
And if you want to be sure it really works, try it somewhere else
Same four letters, a school district instead of a steel mill, and this time the fragile clause isn't a damages cap. It's who pays when a child gets hurt on a piece of playground equipment.
Recital, built by Tanager Compliance, reads facilities and vendor contracts for public school districts: playground equipment maintenance, bus fleet servicing, food service vendors, the contracts a district signs every year without a legal department the size of a steel company's. Branimir Sorrelfield, Cadenwick Unified School District's director of risk and business services, is deciding whether to renew Recital past its pilot year, and he's already skeptical for a plain reason: a school district can't afford to be wrong about who's liable when a child gets hurt.
Different building, same shape of question. A steel mill's damages cap became a playground vendor's liability cap, and the gate logic transferred without changing.
Recital's own eval set is smaller, 95 confirmed disputed facilities and vendor contracts pulled from insurance claims and litigation across public-sector clients, and its measured false negative rate on liability and indemnification language runs a little higher, about 3.1 percent, honestly reflecting the thinner sample. The gate is built the same way, sized to a school district's own budget: any contract worth more than $250,000, any personal-injury liability or indemnification clause, or any insurance-certificate requirement, at any price, routes to Cadenwick's own risk counsel automatically.
The near miss that made Branimir a believer wasn't caught by a decade-old habit the way Osaze's was. It came from a new contracts coordinator three weeks into the job, reading a playground equipment maintenance vendor's contract for the first time with no old habits to fall back on. She asked a plain question nobody senior had thought to ask in years: why did this vendor's liability cap read so differently from the district's usual template. It turned out the cap, buried in a special-conditions page, limited the vendor's exposure to the cost of the equipment itself, nowhere near what a real injury claim could actually run.
The decision Tanager would take back
Recital launched with the same shape of gate Indemnis started with: one blended score, gate anything under a single cutoff. It made sense for a small pilot district with a handful of vendor contracts a year. It stopped making sense the moment Cadenwick's facilities portfolio grew past what one score could safely represent.
Same rank, different lever, mapped straight onto PICK: the position is the same, measured and gated, not claimed perfect. The impact splits the same way, a risk officer who over-restricts loses the whole point of the tool, a coordinator who trusts a wrong clear loses the district's own protection. The cost asymmetry lands on the same kind of clause, just smaller dollars: a routine supply order missed is cheap, a liability cap missed on a facilities contract is the one that can't be undone once a child is actually hurt. And the kill criteria transfer directly: a false negative rate on liability language crossing roughly double today's baseline, or a model swap without a fresh validation report, both mean stop trusting it until it's proven again.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: measured, not perfect, gated by value and clause type regardless of score, and here's the number that would make me pull it back.
Cost: no budget this quarter to widen the eval set. Ship the gate as it stands, labeled honestly as built on a smaller sample, and grow the eval set with every new confirmed dispute instead of waiting to launch until it's larger.
The model got better, for real: say Indemnis's false negative rate drops to 1 percent. The gate doesn't loosen. A lower rate still isn't zero, and the contracts where a miss can't be undone haven't gotten any less serious just because misses got rarer.
Where people run it wrong.
They lead with the accuracy number because it's the one that sounds most impressive, and let a skeptical reviewer poke a hole in the one claim that was never really the point.
They treat a near miss as proof the whole tool is broken, instead of proof of exactly which gate was missing.
They answer a scare by turning the tool off everywhere, instead of narrowing the gate to where the real risk actually lives.
How to use it live. Before naming any number, ask yourself out loud what the reviewer would ask next: "which mistake here can't be undone, and does that one ever reach someone without a person checking it first?" That question alone buys real thinking time, and it's usually the exact distinction the interviewer is listening for.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
PICK: commit to a real position, name who pays for each kind of wrong, find the one that can't be undone, then say what evidence would flip your mind. Built for live "convince me" trust questions.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Zell Kase, who owns Indemnis's trust story at Wrensmoor; Cateline Feathering, Kestrevale Steel's skeptical head of vendor risk; and Osaze Marnier, the contracts manager whose old habit caught Contract 187.
3 · THE POSITION
What's the P step here, in one line?
Tap to flip
ANSWER
We don't promise Indemnis is always right. We promise it's measured, monitored clause by clause, and it can never clear a liability or indemnification section alone once the stakes are high enough to matter.
4 · THE COST ASYMMETRY
Which mistake is cheap, and which one should Cateline actually fear?
Tap to flip
ANSWER
A missed clause on a routine order under $50,000 is cheap, an afternoon and a reread. A missed liability or indemnification clause on a multi-million dollar deal is the one to fear, because it can't be undone once material has shipped.
5 · THE KILL CRITERIA
What evidence would make Zell tell Cateline not to trust Indemnis yet?
Tap to flip
ANSWER
A confirmed false negative rate on liability and indemnification clauses crossing 5 percent, roughly double today's rate, or any model version swap that ships without a fresh 240-contract test run shared before it goes live.
6 · THE OLD DECISION
What decision would Zell take back?
Tap to flip
ANSWER
Scoring the whole contract as one blended confidence number instead of scoring liability-shaped clauses on their own. It made sense for a two week pilot setup. It let a 96 percent document hide one clause that was really a 61.
7 · THE NUMBER
Fill in the blank: Indemnis's eval set has ___ confirmed disputed contracts. It misses a real liability issue about ___ percent of the time.
Tap to flip
ANSWER
240 confirmed disputed contracts. About 2.5 percent, against a 4.1 percent miss rate for one attorney doing a single fast read alone.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which one, and what's the equivalent fragile clause?
Tap to flip
ANSWER
Recital, Tanager Compliance's contract-reading tool for school districts. The equivalent fragile clause is a playground equipment vendor's liability cap, buried in a special-conditions page, far below what a real injury claim could cost.
Check yourself Score: 0 / 0
True or false
1. True or false: gating Indemnis on one blended confidence score for the whole document was a safer design than gating each clause type on its own.
True
False
Show hint
Check what Contract 187's blended score was, and what the buried clause scored on its own.
Show answer
False. A blended score can hide one genuinely risky clause inside an otherwise clean document. Contract 187 scored 96 percent overall while the one clause that mattered would have scored 61 on its own.
Multiple choice
2. Why does the gate route by contract value and clause type, instead of relying on Indemnis's confidence score alone?
A. Because confidence scores are too expensive to calculate for every clause.
B. Because a clause that can't be undone if it's wrong needs a person's eyes regardless of how sure the model sounds, and value and clause type are the two signals that predict which mistakes can't be undone.
C. Because Kestrevale's legal team doesn't trust AI tools in general.
D. Because the confidence score only works on contracts written in English.
Show hint
Look at the C step, cost asymmetry, in the PICK recap.
Show answer
B. A confidence score, even a high one, doesn't tell you whether the mistake it might be making is one you can walk back later. Value and clause type do.
Fill in the blank
3. Indemnis's eval set has ___ confirmed disputed contracts. It misses a real liability or indemnification issue about ___ percent of the time, against ___ percent for one attorney doing a single fast read alone.
Show hint
Check "Let's learn," right after the knowledge spark on false negatives.
Show answer
240 confirmed contracts, 2.5 percent, 4.1 percent. Indemnis already beats a rushed single human read on its own. That's not the reason to trust it. The gate is the reason.
Short answer, name the rejected alternative
4. What alternative did Zell's team consider instead of building separate per-clause gates, and why did it lose?
Show hint
Look at the "three things worth stating directly" paragraph near the end of the PICK recap.
Show answer
Model answer: Simply raising the overall blended confidence bar, say from 90 to 98 percent, instead of building separate gates by clause type. It lost because Contract 187 scored 96 percent blended, above even a stricter overall bar. The flaw was never the height of one number, it was averaging a whole document into one score.
Short answer, apply it yourself
5. Think of an AI tool you've used or heard pitched that made a broad claim like "highly accurate" or "you can trust it." What specific, high-stakes mistake should have gotten its own gate, separate from the tool's overall score?
Show hint
Look for the one kind of mistake in that tool that can't be undone once it happens, not the mistake that's just annoying.
Show answer
Model answer: An AI resume screener pitched as "94 percent accurate." The overall number hid how it performed on candidates with career gaps, a genuinely high-stakes mistake, since wrongly filtering one out can't be undone once they never hear back. That slice needed its own measured rate and its own human check, not a share of one blended average.
Short answer, work the number
6. If the next quarterly batch of confirmed disputes shows Indemnis's false negative rate at 4.8 percent, does that cross the kill line? What about 5.3 percent?
Show hint
Check the kill line chart in the PICK recap and what the K step actually says the bar is.
Show answer
4.8 percent does not cross it. 5.3 percent does. The kill line is 5 percent. 4.8 is close enough to watch closely but still under the line Zell committed to out loud. 5.3 crosses it, and by his own stated criteria, that's the point where he tells Cateline to stop trusting it there.
Before you close the answer
Why this works
Tests whether you can make a reliability case that survives being disbelieved, not just one that sounds confident. A candidate who only reassures is giving a sales pitch. This question is listening for whether your case is falsifiable, and whether the evidence is specific to the model, not generic trust-building.
Follow-up traps
"Why not just review every contract by hand and use Indemnis as a second check instead?" Response: that recreates the exact cost the gate was built to remove, $330 a contract and a two week backlog, while the targeted gate already puts a person on every dollar and clause combination where a real miss would actually hurt.
"Your 2.5 percent number is measured on your own eval set. Isn't that circular?" Response: no. Every one of the 240 contracts is a real, independently confirmed dispute or arbitration outcome pulled from case files, decided by courts and arbitrators, not chosen by Wrensmoor after the fact based on what Indemnis got right.
If pressed
The model version commitment isn't a verbal promise. Any swap of Indemnis's underlying model triggers a full re-run of the 240-contract eval set, and Wrensmoor shares the delta report with the customer's own risk team within five business days. The customer can hold the prior version for up to thirty days if that report shows any drop in performance on the four gated clause types, before the new version ever touches their account.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.