ConceptIntermediateAI Opportunity & Model Strategy / Opportunity identification for AI / #16

Explain how the cost of being wrong should shape which opportunities you pursue first.

BOUND · sequencing Covenance's redline checks, an AI product that reviews and suggests redline edits on legal contracts

Covenance reads a contract next to the version before it and suggests the redline edits a lawyer would normally have to spot by hand. Eustace Sarrazin runs product there. Dovima Beauchene runs client success, and she just got off a call with Severine Verrocchio, senior contracts manager at Northwold Fabrication, who caught a bad redline eleven minutes before she would have signed it. Dovima wants Covenance's next big feature to catch that automatically, every time. Eustace has to work out whether that is actually the right thing to build next.

The direct answer
Rank AI opportunities by how bad it is to be wrong, not by how good it looks to be right. Being wrong has three parts: how bad one wrong output is, how long it takes anyone to notice, and how many other decisions it quietly gets copied into before that happens. Ship the checks where a mistake is caught and fixed in seconds first. Hold the checks where a wrong answer can sit inside a signed contract for months, and spread into other deals, until there is a track record and a real way to test them.
Do this, in order
  1. Rank candidate features by how bad it is to be wrong, not by how impressive it is to be right.Why: this is the one decision the whole answer is built to defend.
  2. Break "cost of being wrong" into three real questions before ranking anything: how bad is one wrong output, how long before anyone catches it, how far can it spread first.Why: a feature can look safe on one of these and dangerous on the other two.
  3. Own the real numbers on both candidates, including how long a wrong answer could sit unnoticed, not just how often the model is wrong.Why: an error rate on its own hides the part that actually does the damage.
  4. Treat cost of being wrong as a range, not a yes-or-no.Why: what makes it worse and what makes it better are both things you can name and check, not a gut feeling.
  5. Sanity-check any ranking built only on upside against the real expected cost of a miss.Why: a catastrophic, slow-to-catch failure can beat a big headline number every time.
  6. Ship the cheap-to-be-wrong feature first, build the eval set on the expensive one while it waits, and gate it behind mandatory review once it ships.Why: sequencing, not refusal, is the actual decision here.

How to answer this, stage by stage

Nobody is grading whether you can recite "break it down, own the numbers, use a range." They are grading whether you can turn "cost of being wrong" into a number that would actually change what a legal-tech team ships next quarter.

01
Scope it to one company, one pitch, one real decision
Say it like this
"Let's ground this. Covenance suggests redline edits on legal contracts. Client success just watched a client almost sign away an uncapped liability claim, and now they want to pitch a feature that catches that automatically, first. That's the specific call I'm going to make: does that ship before or after the boring stuff."
Why this works
Turns a broad prioritization question into one checkable decision instead of a lecture on frameworks.
02
Say the rule out loud before touching a single number
Say it like this
"My rule for sequencing AI features is this: rank by cost of being wrong, not by upside. And cost of being wrong isn't one number. It's three: how bad is a single miss, how long before someone catches it, and how far does it spread before that happens. I'll walk both candidates through all three."
Why this works
Tells the interviewer you have a real test to apply, not a gut call you're about to justify after the fact.
03
Break it down: what "cost of being wrong" actually means, the B step
Say it like this
"One, how bad is a single wrong output, on its own. Two, how long does it take anyone to notice, seconds or months. Three, how far does it spread before it's caught: does it stay stuck in one contract, or does it get copied into the next few deals as a template. Skip any one of those three and you're ranking on vibes."
Why this works
Gives the interviewer a testable structure instead of a single vague word like "risk."
04
Own the real numbers on both candidates, the O step
Say it like this
"Covenance's formatting checker flags something on almost every contract, about 2,640 flags a month at Northwold. A wrong one gets caught in about six seconds, because the lawyer already reads every suggested redline before accepting it. The liability flag only fires on the hard calls, about six contracts a month. If the model gets one of those wrong, nothing draws the lawyer's eye, because the whole point of the feature is telling them not to look twice. It surfaces at the next scheduled audit, about nine months out on average, or later, if a real dispute forces someone to reread the fine print first."
Why this works
A number with a stated source survives a follow-up question. A feeling about which feature matters more doesn't.
05
Show the range, then sanity-check it, the U and N steps
Say it like this
"It's a range, not a switch. A shifted cross-reference sits in the middle, someone usually catches it in a day or two, contained to that one contract. The liability miss is the worst case: undetected for months, and it can get copied into two or three more deals as a template before an audit ever looks. Weight that by how often it actually escalates into a real dispute and you get about $3.3 million a year in expected exposure, against a $450,000 cap the clause was supposed to guarantee. That number is the whole reason you don't ship this the way you shipped the formatting checker."
Why this works
This is the exact check that stops a team from chasing whichever feature would look best in a client demo.
06
Name the direction, reject an alternative, and close on the rule, the D step
Say it like this
"The thing that moves this estimate most isn't how often the model gets the hard call wrong. It's whether we catch a miss before a dispute happens at all. Fix that, and the $3.3 million mostly disappears. So here's the rule: ship the cheap-to-be-wrong feature first. While it's live, start logging every hard liability call next to an attorney's real answer, that's the eval set we don't have yet. We looked at shipping the liability flag first, since it's what almost hurt a client, and dropped it. Reacting to one near miss with an unproven feature is how you build a second, worse near miss. Ship it later, gated behind mandatory sign-off, never auto-cleared."
Why this works
Ends on an operating rule the interviewer can picture actually running, not a promise to be more careful next time.

Let's learn

What happens when an AI product is right almost all the time, and wrong in exactly the way nobody is watching for?

Say a company builds a tool that reads a legal contract next to the last signed version and suggests the edits, the redlines, a lawyer would normally have to spot by hand. Before the tool existed, a senior contracts manager read every clause herself: checking definitions, checking cross-references, checking every non-standard term against what her company usually accepts. About two hours for an average supplier contract.

With the AI's redline suggestions doing the first pass, that same review drops to about twenty-five minutes for most contracts.

Hand sketched labeled parts diagram titled What cost of being wrong actually breaks into. A scale icon at the center labeled Cost of being wrong, with three callouts around it reading How bad, once, How slow to catch, and How far it spreads.
Three separate questions hiding inside one phrase. Miss any one of them and you are ranking features on a feeling.

Here is the turn. The extra minutes saved are not really the point. The tool does not make one kind of mistake, it makes two very different kinds, and only one of them looks like a mistake at all. Some wrong flags are loud and useless: a defined term flagged as inconsistent when it isn't, easy to catch and dismiss on the spot. Others are quiet. They tell the lawyer a clause is fine when it isn't, and quiet is exactly what a busy person stops double-checking.

Knowledge spark: what's a false negative? The model saying nothing is wrong when something actually is. It's the quiet kind of wrong. Nothing on the screen tells anyone to look twice, so nobody does, until the thing it missed matters on its own.
A wrong flag that gets caught in six seconds and a wrong flag that hides for nine months are not the same size of wrong.

What it costs at its worst: a team builds the feature that catches rare, severe clauses, trusts it because it demoed well, and a wrong "no issue" call sits inside a signed contract for months. Worse, that contract becomes the template for the next deal, and the next one after that, so one quiet miss ends up baked into three live contracts before anyone rereads the fine print. If a real dispute lands before anyone catches it, the company is arguing over uncapped damages instead of a number it planned for.

The choice I would take back Covenance's roadmap review let whichever incident scared the room most decide what got built next, not how expensive or slow a miss would be to catch. That was fine back when every candidate feature was roughly as safe to get wrong as the last one. It stopped being fine the moment "scariest in the room" quietly stopped matching "cheapest to get wrong."

What I would leave alone: the formatting checker doesn't need any of this caution. Every suggestion it makes gets a human glance before it's ever accepted, that's just how the review screen works, and a rejected flag changes nothing downstream. Slowing that feature down with extra review gates would be solving a problem it doesn't have.

The lesson: a feature that gets applause in a pitch meeting and a feature that's actually safe to ship first are answering two different questions. One asks what would impress a client. The other asks what happens the day it's wrong. Answer the second one first.

Now here is the same thing as a story

The short version is above, for saying out loud in an interview. Read this one for the actual Tuesday the near miss and the numbers ended up on the same desk.

Eustace Sarrazin has a habit his old boss used to tease him about: before he'll agree to build anything, he wants the sentence that describes exactly how it fails. He'd run product at Covenance for three years, and the habit had saved the team more than once, usually by killing a feature nobody wanted to admit was half-formed.

Dovima Beauchene is good at a different thing: a room. She runs client success, and she'd just come off a call with Severine Verrocchio, senior contracts manager at Northwold Fabrication, an industrial parts manufacturer that ran about 240 supplier contracts a month through Covenance.

Hand sketched two panel comparison titled Covenance's two candidate redline checks. Left panel, a document icon, labeled Formatting check, caption reads Every contract, a wrong flag is caught in seconds. Right panel, a scale icon, labeled Liability flag, caption reads One in forty contracts, a wrong call can hide for months.
The two candidates as Dovima first laid them out, before anyone had run a single number against either one.

Severine had called Dovima that morning, still a little shaken. A supplier's redline had come back on a routine renewal, every field marked clean by Covenance's checker. Eleven minutes before the signature deadline, she reread clause 9 anyway, an old habit from before the tool, not because anything on the screen told her to. The redline had quietly stripped the carve-out that kept a gross-negligence claim inside Northwold's standard $450,000 liability cap. Nothing had flagged it. She caught it by instinct, with eleven minutes to spare.

Hand sketched horizontal timeline titled Eleven minutes at Northwold. Five milestones: Redline returned, every field marked clean. Queued to sign, 11 minutes left. Severine rereads clause 9, old habit not a flag, this milestone emphasized in rust orange. Carve-out is gone, gross negligence uncapped. Signature held, caught by luck not by design.
Nothing on Severine's screen changed color that morning. The only thing that caught it was a habit the tool was never designed to need.

Dovima pitched it at the next roadmap review as Covenance's obvious next feature: teach the model to flag exactly this, automatically, before it ever reaches a lawyer's desk. The room liked it. A near miss with real names attached is a better story than a spreadsheet.

Eustace was asked to scope it. He started, as he always did, by asking what "wrong" would cost, not just how often it would happen. The formatting checker was easy: Northwold's lawyers read every suggested redline before accepting it, so a bad flag gets a six-second glance and a dismissal, out of about 2,640 flags a month. Nothing propagates from a rejected flag. Worst case, it costs the lawyer a handful of seconds, over and over, and never once leaves the room.

The liability flag was a different shape of problem entirely. Hard calls like Severine's show up in about one in forty contracts, roughly six a month at Northwold. On Covenance's own held-out set of past disputes, the model got calls that hard wrong about one time in twelve. Say yes to a feature that quietly clears a clause instead of flagging it, and that miss doesn't announce itself. It surfaces at Northwold's next scheduled audit, every eighteen months, or later, if a real claim forces someone to reread the contract first.

Knowledge spark: what's a held-out set? A pile of real past cases kept aside and never shown to the model while it's being built or tuned. Testing against those cases is the only fair way to know how the model does on something it hasn't already seen the answer to.

Eustace ran the arithmetic all the way through. Five times out of six, a wrong call gets caught at the audit, costing about $6,000 to fix per contract, and by then the flawed template has usually been copied into two more deals, so call it $18,000. One time out of six, a real dispute lands first, and outside counsel's own estimate on Severine's near miss put the uncapped exposure at up to $3,200,000. Weight those two outcomes by how often each one actually happens, and the expected cost of one liability miss lands at about $548,000. Six misses a year, at Northwold alone, comes out to about $3.3 million a year in expected exposure.

Hand sketched decision tree titled How Covenance almost ranked its roadmap. Root box reads Which redline check ships first, branching to four outcomes: ranked by upside alone leads to the flag feature no eval set. Ranked by demo excitement leads to the near miss retold as a pitch. Ranked by cost of being wrong leads to formatting first flag gated, this outcome outlined in rust orange. Ranked by both sanity checked leads to the choice that actually held.
Four ways the same roadmap meeting could have gone. Only one of them survives the actual arithmetic.

He brought the number back to the room, next to the $450,000 cap the clause was supposed to guarantee. "We're not saying no to the liability flag," Eustace told Dovima. "We're saying not yet, and not unsupervised. Right now, shipping it with the same trust as the formatting checker would be worth about $3.3 million a year in exposure we can't see coming."

Dovima looked at the number longer than anyone else in the room. She hadn't been wrong that Severine's near miss mattered. She had just been ranking by how loud the story was, not by how the two features actually failed.

Covenance shipped the formatting checker's next version that quarter, on schedule. The liability flag stayed on the roadmap, rebuilt as a tool that surfaces the clause, the redline, and a first-pass read to a human reviewer, who has to sign off before anything is marked clear. Every one of those sign-offs gets logged, right or wrong, building the eval set nobody had six months earlier.

What Eustace would tell his past self: a near miss makes you want to fix the exact thing that almost went wrong, right now, with whatever's fastest. That instinct is honest, and it is also how you ship the next unproven thing straight into production with no way to check it. The six seconds nobody worried about were never the risk. The nine months nobody was counting were.

BOUND, for pricing what a quiet miss at Covenance actually costs

Not a way to make an exciting pitch sound suspicious. BOUND turns "this feels risky" into a number Eustace, or anyone else at Covenance, could actually defend in the room.

BBreak it down. What does "cost of being wrong" actually split into?
Three separate questions, not one. How bad is a single wrong output, by itself. How expensive and slow is it to detect, seconds or months. How far can it spread before anyone catches it, contained to one contract, or copied into the next few as a template. A feature can pass the first question and fail the other two, so all three have to be checked, not just the loudest one.
Skip this split and "cost of being wrong" stays a feeling instead of three checkable questions.
Hand sketched two axis quadrant diagram titled Where three real redline misses actually land. Horizontal axis How fast someone catches it, from instantly to months later. Vertical axis What it costs if wrong, from almost nothing to a real claim. Three labeled dots: Formatting flag sits lower left, caught instantly and costs almost nothing. Section cross reference sits near the middle, a moderate case. Liability clause sits upper right, caught months later and costs a real claim.
Cost of being wrong is not a switch. It is two dials moving together: how fast someone catches it, and what it costs while they haven't.
OOwn the numbers. Where does each one actually come from?
Formatting checker: about 2,640 flags a month at Northwold, roughly 4 percent wrong, caught in about 6 seconds each because the lawyer reads every suggestion anyway. Worst case, about $244 a year in lawyer time, company-wide at Northwold. Liability flag: about 6 hard calls a month, wrong about 1 time in 12 on Covenance's held-out set of past disputes, so about 6 misses a year. Five of six get caught at an 18-month audit for about $18,000 each once propagation is counted. One of six escalates into a real dispute first, at up to $3,200,000, based on outside counsel's own estimate on Severine's near miss.
A number only counts as owned if you can say exactly where it came from when someone pushes on it in the room.
Knowledge spark: what's an eval set? A pile of real examples with a known right answer, used to check how often a model actually gets something right before anyone trusts it. No agreed right answer yet means no eval set, and no eval set means nobody can honestly say the model is ready.
What one liability-clause miss actually costs Northwold in a year, expected
$0 $1.0M $2.0M $3.0M Northwold's per-contract cap, $450,000 $3.29M / year, expected Liability-flag miss, weighted Formatting flag, about $244/yr (too small to draw here)
Caught at the 18-month audit, 5 years out of 6Caught only after a real dispute, 1 year out of 6
The audit-only segment alone, $90,000 a year, would already be a real number. The dispute segment is what turns it into a company-changing one.
UUse a range, not a binary.
Three real points, not two. Formatting flags: caught instantly, costs almost nothing, contained to nothing because nothing gets accepted unseen. Section cross-references: a redline shifts clause 4.2 to 4.3, but a reference elsewhere still points to 4.2, usually caught within a day or two when someone assembles the signature packet, contained to that one contract. Liability clauses: caught in months, not days, and able to spread into two or three more contracts as a reused template before anyone rereads the fine print. What makes it worse: higher stakes per instance, slower detection, more downstream spread. What makes it better: lower stakes, fast feedback, a blast radius that stays contained.
A single "how risky is this" answer here repeats Dovima's exact mistake, treating three very different failure shapes as if they were one.
NNail the sanity check. Does ranking by upside alone survive contact with the real numbers?
A team that ranks candidate features purely by potential upside, "this would impress every client," is only looking at half the picture. The liability flag has the bigger headline upside. It also has an expected cost of about $3.3 million a year if shipped with no guardrails, more than seven times the $450,000 cap the missing clause was supposed to guarantee. A modest, boring feature that is safe to get wrong can beat a flashy one with a catastrophic, slow-to-catch failure mode, every time you actually run the numbers instead of trusting the pitch.
This is the exact check that would have stopped the pitch before it shipped an unproven feature straight into production.
What would shrink that $3.3 million most
Always catch it before a dispute -$3.18M Cut the wrong-call rate in half -$1.65M Stop template reuse entirely -$0.06M
Detection pathModel wrong-call ratePropagation count
Fixing detection speed does fifty times more work than stopping propagation. That is the single fact that decides what Covenance builds next.
DDirection. Which assumption moves this most, and what's the actual decision?
Not the model's wrong-call rate, and not how many contracts a template gets copied into. The single biggest lever is whether a miss gets caught before it ever escalates into a real dispute, the chart above shows that one fix alone erases most of the $3.3 million. So the real decision is a sequencing rule: ship the cheap-to-be-wrong feature first, always. Build the eval set on the expensive one while it's still on the shelf. Ship it later, behind mandatory attorney sign-off with no auto-clear, so a wrong "no issue" call never reaches a signature unseen.
Naming the fact that actually swings the estimate, instead of the biggest headline number in the room, is what separates a real estimator from a confident guesser.
Hand sketched left to right flow diagram titled The order Covenance actually ships in. Five connected boxes reading Formatting, Log calls, Eval set, this box outlined and emphasized in amber, Gated flag, Loosen gate.
The liability flag never gets cancelled. It just has to wait its turn behind a real eval set.

One alternative considered and rejected: building the liability flag first, right after Severine's near miss, and shipping it fast because it was what almost hurt a real client. It lost, because Covenance had no eval set for it yet, no held-out sample big enough to say how often it would actually be right, and shipping an unproven autonomous flag into the exact spot Severine's own instinct had just saved them from would have traded one near miss for the risk of a real one. The AI-specific failure mode underneath all of this is quiet and easy to miss: a model can say "standard, no issue" on a rare, severe clause with the exact same confident tone it uses on a clause that really is standard, and nothing about the interface tells a reader to look twice. The guardrail is mandatory human sign-off on every hard liability call, with no auto-clear, plus logging every one of those calls against the real redline, which is what actually builds the eval set. And the trade-off is real, not free: routing every hard call through a person costs turnaround time and senior counsel hours that the formatting checker never spends, trading speed and cost for a wrong call that can never reach a signature unseen.

And if you want to be sure it really works, try it somewhere else

Same five letters, a radiology reading room instead of a legal desk, and this time the tempting feature is catching a rare finding instead of catching a rare clause.

Umbergate gives a second read on chest X-rays before a radiologist signs off, and it runs the same choice Eustace did, on a different kind of file. One candidate check flags poor image quality, bad positioning, motion blur, before a radiologist ever reads the scan. The other flags a subtle, rare finding, a small pneumothorax, a nodule easy to miss on a busy shift.

Run BOUND on it. Break it down: the quality flag happens on every scan, and a wrong flag gets caught before the patient even leaves the room, the technologist retakes it or waves it through in under a minute. The rare-finding flag happens on maybe 1 in 150 scans, and a miss doesn't get caught by the normal read, because the whole point of the flag is telling the radiologist not to look twice. Own the numbers: at a hospital reading about 3,000 chest films a month, quality misfires run about 2 percent, 60 a month, each one a 60-second retake decision made by a person, not the model. Rare-finding hard calls show up on about 20 scans a month, wrong about 1 time in 10 on Umbergate's own held-out set, so about 2 misses a month. Without a forced follow-up, the average time before one surfaces on its own, a return visit, a later scan, runs about 5 weeks, and sometimes it never surfaces at all.

Hand sketched numbered icon list titled Why Umbergate's scan quality flag is safe to ship first. Four rows: one, a gauge icon, A bad flag is seen before the patient leaves the room. Two, a person icon, The tech, not the model, makes the retake call. Three, a document icon, Nothing about the read gets written down yet. Four, a scale icon, A missed rare finding waits for symptoms, not a glance.
Four separate reasons the quality flag can absorb a wrong guess, and none of them apply to the rare-finding flag.
Where Umbergate's answer genuinely differs A missed rare finding can cost a delayed diagnosis, not a dollar figure, so the arithmetic can't run the same expected-value math Northwold's contracts allow. But the shape underneath is identical: the cheap-to-be-wrong check goes first, the expensive one waits for a forced follow-up habit and a real eval set, and nothing rare and severe ever gets auto-cleared without a second set of eyes.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Name the three parts of cost of being wrong, own one real number on each candidate, and say plainly that a range beats a single confident guess.
Cost: no time to build the full weighted estimate before the pitch. Shrink it to the one comparison that would actually decide the argument, detection time, and say honestly that the rest is a rough range.
The model got better, for real: a new release cuts the wrong-call rate on hard liability clauses in half. Recheck the sequencing anyway, since the chart above shows the error rate was never the biggest lever, detection speed was.

Where people run it wrong.
They rank by how big a single case could be instead of by how fast anyone would notice being wrong.
They treat "the model is usually right" as proof it's safe to ship unsupervised, instead of asking what happens on the rare time it's confidently wrong.
They react to one near miss by rushing the exact feature that almost caused it, with no eval set to check it against.

How to use it live. Before answering, ask yourself one plain question out loud: "if this feature is wrong, how fast would anyone actually know?" If the honest answer is "not for months," say that out loud, and let it be the reason the feature waits, not a detail you hope nobody asks about.

Flashcards (tap any card to flip it)

1 · THE METHOD
What method fits deciding which of two AI features to ship first, when the real question is how bad it is to be wrong?
Tap to flip
ANSWER
BOUND: break it down, own the numbers, use a range, nail the sanity check, name the direction. Built for turning "this feels risky" into arithmetic you can defend.
2 · WHO'S IN IT
Who is this answer about?
Tap to flip
ANSWER
Eustace Sarrazin, Covenance's product lead. Dovima Beauchene, Covenance's client success lead. Severine Verrocchio, senior contracts manager at Northwold Fabrication, who caught the near miss.
3 · THE BREAKDOWN
What three questions does "cost of being wrong" actually split into?
Tap to flip
ANSWER
How bad is one wrong output, on its own. How long before anyone notices it. How far can it spread, into how many other contracts or decisions, before that happens.
4 · THE TWO NUMBERS
Fill in the blank: a wrong formatting flag gets caught in about ___. A wrong liability call can sit uncaught for about ___.
Tap to flip
ANSWER
About 6 seconds. About 9 months, the midpoint of an 18-month audit cycle, and that's the fast path, since a real dispute can surface it later, or never.
5 · THE RANGE
Name the middle case between "caught instantly" and "hidden for months," and where it lands.
Tap to flip
ANSWER
A shifted section cross-reference. Usually caught within a day or two when someone assembles the signature packet, contained to that one contract, no real dollar exposure.
6 · THE SANITY CHECK
Why was Dovima's pitch to build the liability flag first the wrong call, even though the underlying near miss was real?
Tap to flip
ANSWER
She ranked by how scary the story was, not by the actual expected cost. About $3.3 million a year, expected, against a $450,000 cap the missing clause was supposed to guarantee.
7 · THE DIRECTION
Which single assumption swings the $3.3 million estimate the most?
Tap to flip
ANSWER
Whether a miss is caught before it ever escalates into a real dispute, not the model's wrong-call rate and not how far a template spreads. Fixing detection alone erases about $3.18 million of the $3.3 million.
8 · SAME METHOD ELSEWHERE
Section 4 runs BOUND again on a different product. Which one, and what's the same shape?
Tap to flip
ANSWER
Umbergate, a second-read radiology tool. Same shape: a scan-quality flag, caught in a minute by the technologist, ships before a rare-finding flag, which can go unnoticed for weeks without a forced follow-up.

Check yourself Score: 0 / 0

Multiple choice
1. Why does the liability-clause flag deserve more caution than the formatting checker, even though the formatting checker is wrong far more often in raw count?
  • A. A liability miss costs far more per instance, takes months to catch, and can spread into other contracts before anyone notices.
  • B. Covenance's engineers don't have the skills to build the liability flag yet.
  • C. Northwold refused to let Covenance touch liability clauses at all.
  • D. The model can't process liability language, only formatting.
Show hint
Check the B and O steps.
Show answer
A. High per-instance stakes, slow detection, and real propagation are the three things that make a miss expensive, not the raw count of how often it happens.
Fill in the blank
2. A wrong formatting flag at Northwold gets caught in about ___. A wrong liability-clause call can sit uncaught for about ___ before the next scheduled audit.
Show hint
Check the O step and the story's numbers.
Show answer
6 seconds. 9 months. The formatting flag gets a human glance by design. The liability flag's whole job is telling the lawyer not to look twice, which is exactly why a wrong one hides.
True or false
3. True or false: stopping a flawed contract from ever being copied into a new deal as a template is the single biggest lever for shrinking Northwold's $3.3 million expected annual cost.
  • True
  • False
Show hint
Check the sensitivity chart in the N and D steps.
Show answer
False. Stopping propagation only shrinks the estimate by about $60,000. Catching every miss before a real dispute happens shrinks it by about $3.18 million, by far the biggest lever.
Short answer, where it would not matter
4. Name a place in Covenance's product where this extra caution would NOT apply, and say why.
Show hint
Check "what I would leave alone" in Let's learn.
Show answer
Model answer: The formatting checker. Every suggestion already gets a human glance before it's accepted, and a rejected flag changes nothing downstream, so there's no hidden failure mode to guard against.
Short answer, apply it yourself
5. Think of an AI feature you use or could build yourself. Name one place a wrong output would be obvious in seconds, and one place it could hide for weeks.
Show hint
Ask who would notice the mistake, and how long it would take them.
Show answer
Model answer: An AI that drafts email replies: a wrong tone is obvious the second you reread the draft before sending. An AI that auto-categorizes expense receipts for tax filing could misclassify a rare deductible category and nobody notices until an audit, months later.
Short answer, work the number
6. If Covenance's wrong-call rate on hard liability clauses doubled, from about 6 misses a year to 12, but every one of those misses got caught at the audit before any real dispute, would the expected annual cost go up or down from today's $3.3 million? Show the arithmetic.
Show hint
Use the audit-only cost of $18,000 per miss, with no dispute-path risk.
Show answer
Down, sharply. 12 misses times $18,000 equals $216,000 a year, far below today's $3.3 million, because removing the dispute-path risk matters more than doubling the error rate ever could.
Before you close the answer
Why this works
Tests whether a candidate can resist the pull of a scary, recent near miss and instead sequence a roadmap by real risk-adjusted arithmetic, when the wrong call would ship an unproven feature straight into the exact spot that almost caused real harm.
Follow-up traps
"Isn't $3.3 million an exaggerated worst case, since only 1 in 6 misses actually escalates?" Response: it's an expected value, not a worst case, that's exactly what "weighted by 1 in 6" means, and even the audit-only floor, $108,000 a year, still dwarfs the formatting checker's cost.

"If the liability feature is this risky, why build it at all?" Response: because sequencing, not refusal, is the answer. Ship it once there's a real eval set and mandatory sign-off, that's a guardrail, not a ban.
If pressed
The 1-in-12 wrong-call rate came from scoring the model against 340 real historic liability determinations Northwold's own outside counsel had already resolved, held out of any prompt tuning. That sample is still too small to trust the flag unsupervised, which is exactly why the gate stays on until it grows.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more