CaseIntermediateModel Fluency & the AI PM Role / AI PM vs traditional PM vs technical PM / #4
How does the discovery phase differ when feasibility is genuinely unknown until you build?
BOUND · a legal research AI that finds and checks case citations for lawyers
Holdfast is Falconbridge's AI tool for litigation teams: it reads a lawyer's draft argument and finds real case law to back it up, citation and all. Vianne Underhill owns its rollout. Fairhurst and Merriwether, the mid-size litigation firm running the pilot, will not put Holdfast anywhere near a real filing until someone can tell managing partner Zdenka Woolgar exactly how it fails, and how anyone would know.
The direct answer
Run discovery as one timeboxed spike, not a written spec. Build the narrow slice that finds and verifies case citations, test it against a real labeled sample, and come back with a number instead of an opinion. On 150 real citations, 23 came back wrong, and the layer meant to catch the dangerous kind, a real case cited for a point it never actually decided, only catches somewhere between 64 and 86 percent of them. That range, not the overall 84.7 percent that looks clean, is the number that decides whether Falconbridge builds this, narrows it, or kills it.
Do this, in order
Run discovery as one timeboxed spike that hands back a number, not a document.Why: a spec with a made-up accuracy target is a guess wearing a requirement's clothes, and nobody has seen what Holdfast actually gets wrong yet.
Score every citation into its real flaw type, not one pass or fail.Why: fabricated and stale citations are basically solved by a database lookup, and blending them with mischaracterized ones hides which flaw actually decides the answer.
Give the verification layer's catch rate as a range, not a point.Why: it's measured on a small holdout by a model doing the catching, and a single tidy number would hide how much that estimate could move.
Check the result against what a filed brief can actually survive, not against what sounds impressive.Why: 84.7 percent clean sounds like a passing grade until you remember the other 15.3 percent includes citations that could get a lawyer sanctioned.
Name the one number that decides proceed, pivot, or kill: the mischaracterization catch rate.Why: existence and staleness are already solved by a lookup, so this is the only number still standing between a working product and a dangerous one.
Reject shipping straight to a live pilot to see what happens.Why: a fabricated citation reaching one real filing isn't a bug report, it's a career, and that's not a place to learn by doing.
How to answer this, stage by stage
Nobody is grading whether you can say "we'd run a spike" and sound careful. They're grading whether you can turn a genuine unknown into a real number, and a decision that survives someone asking where that number came from.
1
Scope it to one real go or no-go decision
Say it like this
"Let me scope this. Holdfast is Falconbridge's tool that reads a lawyer's draft argument and finds real case law to back it up, citations included. Vianne Underhill owns the rollout. Fairhurst and Merriwether, the pilot firm, need one thing before they'll trust it on a real filing: can Holdfast's citations be trusted, and how would anyone actually know. I'll answer against that."
Why this works
One real product and one real decision keeps the answer from drifting into a general essay about AI in law.
2
Say what discovery actually has to produce, before touching a number
Say it like this
"Normal discovery here would mean interviewing litigation partners and writing a spec that says citations must be, say, 98 percent accurate. I'm not doing that, because nobody in that room has seen what Holdfast actually gets wrong yet. Discovery has to be a working slice, tested against real citations, timeboxed, that comes back with a number."
Why this works
This is the reframe the whole answer turns on. No spec can substitute for testing something nobody has built yet.
3
Own every number in the spike's design
Say it like this
"Here's the spike. 150 real citations, pulled from actual brief sections across five practice areas, 30 each. Three weeks. One engineer and one contract attorney at half time, scoring each citation as clean, fabricated, mischaracterized, or overruled. Three weeks builds the narrow retrieve-and-verify slice without it turning into the whole product. 150 is enough to see all three flaw types show up more than once in every practice area."
Why this works
Every number has a reason attached, so nobody can ask "why 150" and get a shrug back.
4
Report the raw result by flaw type, not one blended score
Say it like this
"Of the 150, 127 came back clean. 6 were fabricated, cases that don't exist. 3 were real cases that had actually been overruled. And 14 were real cases that simply didn't say what the brief claimed they said. That last group is the one that should worry you."
Why this works
A single "84.7 percent accurate" would have buried the one number that actually matters inside a number that sounds fine.
5
Give the residual as a range, not a guarantee
Say it like this
"We built a verification layer to catch the bad ones before a lawyer ever sees them. Existence and staleness are basically solved, that's a database lookup, and it caught all 9 of those. Mischaracterization needs judgment, and on our holdout it caught somewhere between 9 and 12 of the 14. So somewhere between 2 and 5 citations out of every 150 would reach a lawyer's screen looking clean when they weren't."
Why this works
A range is the honest answer here. A single number would claim more certainty than a small holdout can actually back up.
6
Nail the sanity check against what a filing can survive
Say it like this
"Is 2 to 5 in 150 good enough? Compare it to two things. Courts have already sanctioned lawyers for filing briefs with fake AI-generated citations, so the bar here isn't 'better than nothing,' it's 'safe enough that a lawyer's normal read-through catches the rest.' And Holdfast's raw mischaracterization rate, 9.3 percent, is actually worse than a rushed associate checking their own citations alone, which runs about 5 percent. The model isn't already better than a careful human. The verification layer is the only reason this is even close."
Why this works
This is the hardest step, and it's what turns "23 out of 150 were wrong" into a real decision instead of a number nobody acts on.
7
Name the one number that decides it, and close
Say it like this
"So here's the call. The mischaracterization catch rate, 64 to 86 percent, is the number that decides this, not the overall 84.7 percent. If a bigger validation holds that above 80, we build the full generate-and-cite feature with mandatory flagging. If it lands in the middle, we ship a narrower verify-only version first. If it can't clear roughly 50, we kill the generate-and-cite idea and route every citation to a lawyer, full stop."
Why this works
It leaves the interviewer with a decision they could check later against real evidence, not a feeling.
Let's learn
Holdfast reads a lawyer's draft argument and finds real court cases to back it up. It writes the citation, quotes the part that matters, and hands it back ready to drop into a brief.
A citation looks like one small fact. It's really four separate things that all have to hold at once, and only two of them are a simple lookup.
Before Holdfast, an associate at Fairhurst and Merriwether spent four to six hours on a single motion just finding and checking case law: searching a database, reading each case's actual holding, and confirming nothing had been overruled since. With Holdfast's first draft in hand, that search time drops to under twenty minutes. The associate reviews what Holdfast found instead of starting from a blank search box.
Five steps. The fourth one, verifying what the model just wrote, is the step the whole feasibility question actually turns on.
Knowledge spark: what's a mischaracterized citation?
A real case, cited for a point it never actually decided. Not a typo, not a broken link. The case exists, the page number is right, and it still says the opposite of what the brief claims. Courts have already sanctioned lawyers for filing briefs full of fake, AI-invented cases, more than once. This is the same failure with the citation itself real.
Falconbridge ran a three week spike before promising anyone anything. 150 real citations, pulled from five practice areas at Fairhurst and Merriwether, scored one at a time by Jerric Stanmore, a contract attorney hired just for this. 127 came back clean. 6 cases didn't exist at all. 3 were real but had been overruled. And 14 were real cases that simply didn't say what the brief claimed.
The 150-citation spike, broken into clean and three flaw types
CleanFabricatedMischaracterizedOverruled, not flagged
84.7 percent clean sounds like a passing grade. It hides that only 6 plus 3 of the wrong ones are a simple lookup away from solved. The 14 mischaracterized ones are the real question this spike exists to answer.
Three separate tripwires. A citation only clears on its own when none of the three fire, and the middle one is the only tripwire that isn't a certainty.
The 23 wrong citations were never the real problem. The real problem was whether anything would catch the 14 that needed judgment before a lawyer trusted them.
A fabricated or badly mischaracterized citation that reaches a filed brief doesn't just embarrass someone. It can get a lawyer sanctioned, fined, and named in a public order any client can read. Fairhurst and Merriwether weren't going to risk that to save an associate four hours.
The choice I would take back
The first version of the spike just marked each citation right or wrong, one column, to move fast inside the three week window. That hid the fact that fabricated and stale citations were basically solved by a simple lookup, while mischaracterized ones were the only real unknown. I'd split the scoring into flaw types from day one.
What I would leave alone: formatting. Whether a citation follows the firm's house style, commas in the right place, doesn't need a feasibility spike. That's ordinary string matching, already reliable, and testing it here would have burned time on the one thing that was never actually in question.
The lesson: a feasibility spike isn't a smaller version of the product. It's the one question that decides whether to build the product at all, answered with real numbers instead of a guess that happens to sound confident.
Now here is the same thing as a story
The short version above is what you actually say in the room. Read this one when you want to feel why a single citation, almost logged clean in four seconds, nearly decided the whole answer.
Jerric Stanmore has checked citations for eleven years, first as a summer associate, now as the person firms call when they need someone fast and exact. He can read one paragraph of a court's holding and tell you in ten seconds whether it says what a brief claims it says. He has a habit nobody taught him: he always reads the actual paragraph a pin cite points to, never just the case name and the page number, because the two can drift apart in ways a fast scan won't catch.
Falconbridge brought him in for three weeks, at half time, to score Holdfast's citations one at a time. Clean, fabricated, mischaracterized, overruled. Real cases, pulled from real work Fairhurst and Merriwether had already filed.
The first week went well. Contract and employment citations mostly checked out, and the ones that didn't were easy calls: a case that plainly didn't exist, a holding that had obviously been reversed. Jerric started to relax into the rhythm of it, the way anyone does with a task that keeps agreeing with them.
Then insurance coverage.
Citation 94 was a coverage dispute, cited for the idea that a specific policy exclusion didn't apply to a particular loss. It was formatted perfectly. Holdfast's own confidence score on it read 91 out of 100. Jerric almost logged it clean on the strength of the case name and the pin cite alone, the way he might have by citation 200 if the rhythm had kept holding.
He read the paragraph anyway. Old habit.
The case didn't just fail to support the point. It said the opposite. The court had specifically held that the exclusion did apply, on nearly identical facts. Cited the way Holdfast had cited it, the case would have told a judge the reverse of what the actual court decided.
The whole spike, in one picture. Existing is a fact you can check in seconds. Being right is a judgment call, and citation 94 passed the first test and failed the second.
We didn't almost ship a typo. We almost shipped the exact opposite of what a real court said, with a case name and a page number attached to make it look checked.
Jerric flagged it. It took him four minutes to catch, on a citation Holdfast had scored 91 out of 100 sure of itself.
Here's the decision that traces back to a kickoff call, ten days before the spike started. Someone had asked whether the scoring should track flaw type separately from day one, or just mark each citation right or wrong to move faster inside three weeks. The team picked one column. Reasonable, given the calendar. Splitting into flaw types meant more setup, more categories for Jerric to keep straight while moving fast.
I would take that back. Not because one column is wrong exactly, it is faster to build. But by day 9, with one column, all we would have had was "84.7 percent clean," a number that sounds like a passing grade. We would have had no way to tell Zdenka Woolgar the thing buried inside that 15.3 percent: that only 6 of the wrong ones were fabricated, easy to catch, and only 3 were stale, also easy to catch. The other 14, the mischaracterized ones, are the only kind a person like Jerric can catch and a lookup can't. That's the entire question this spike exists to answer, and one column would have hidden it inside an average.
Run the spike again with flaw types split from day one, the way it actually ran. By day 12, nine days early, Vianne already had the number that mattered: a verification layer catching mischaracterization somewhere between 64 and 86 percent of the time, tested against Jerric's own scoring. She didn't need all three weeks to know what to bring Zdenka. She needed the right four columns.
One design tells you whether Holdfast is roughly right. The other tells you exactly which kind of wrong could end a career, and whether anything catches it before a judge does.
What I would tell myself, back on that kickoff call: a single column isn't really faster. It just moves the cost from this week to the week someone asks why the number everyone trusted didn't say what it actually meant.
Three weeks after that call, Vianne told Zdenka Woolgar the truth in one number, not a pitch. Build it, but only as verify-first, until the catch rate proves out past 80. Fairhurst and Merriwether signed on for that version. Nobody ever found out about citation 94 from a judge.
BOUND, or turning "nobody knows yet" into a number Zdenka could act on
Not a way to sound more careful about an unknown. BOUND is what turns "we can't know until we build it" into a real number, and a decision that survives someone asking exactly where that number came from.
BBreak it down. What does the discovery question actually need?
"Can we trust Holdfast's citations" isn't one fact, it's a question hiding four things: how many citations a lawyer actually relies on, what share of those are wrong and in which of three ways, how much of the dangerous kind a verification layer catches, and what's left over once it doesn't. Say the shape before naming a single figure, or discovery quietly becomes an opinion dressed as a plan.
Skip this step and every number after it is a guess wearing a decision's clothes.
OOwn the numbers. Where did each one come from?
150 real citations, pulled from actual brief sections across five practice areas at Fairhurst and Merriwether, 30 each, so no single practice area could hide inside the average. Three weeks and one contract attorney at half time, because that's what the budget could bear before either wasting runway or shipping blind. 6 fabricated, 14 mischaracterized, 3 overruled, scored by Jerric Stanmore against the actual holding text, not just the case name.
Owning a number means saying where it came from and why that size, not just stating a figure that sounds specific.
UUse a range, not one point.
The verification layer's semantic check isn't a certainty either. On the 14 mischaracterized citations in the holdout, it correctly flagged somewhere between 9 and 12 of them, an estimate from a small sample, not a census. Push that range through the math: somewhere between 2 and 5 mischaracterized citations out of every 150 reach a lawyer's screen looking clean when they aren't. A flat "about 3.5" looks tidier and hides how much this could move with more data.
This is the direct answer's arithmetic, in one step. The range is the honest version. The point estimate is the comfortable one.
Mischaracterized citations reaching a lawyer unflagged, as the verifier's catch rate moves
Mischaracterized citations reaching a lawyer unflagged, out of 14Our two measured endpoints
The line carries the real arithmetic: residual equals 14 times one minus the catch rate. Our measured range sits mostly left of the 80 percent line the team set for shipping widely, which is exactly why the decision is pivot, not proceed, until a bigger validation moves it.
NNail the sanity check. Does it survive a real comparison?
Two comparisons, and they both point the same way. Courts have sanctioned lawyers for filing briefs with fake, AI-invented citations, so the bar here was never "better than nothing," it's "safe enough that a lawyer's ordinary read-through catches whatever slips past verification." And Holdfast's raw mischaracterization rate, 9.3 percent, is worse than a rushed associate checking their own citations alone, measured internally at about 5 percent. The model isn't already beating a careful human on this narrow task. The verification layer is the only reason the product is even close to viable.
The hardest step, and the one that turns a number into a reason to act instead of a number to note in a status update.
All four numbers on the same line. The measured range straddles the ship-it line instead of clearing it, which is the whole reason this became a pivot decision and not a launch.
DDirection. Which assumption would move it most?
Not the sample size or the practice-area split, those are steady and barely move the answer if they're off by a little. It's the verification layer's own catch rate on mischaracterization. That 64-to-86 range comes from testing a small model against a small holdout, and it's the single biggest lever in the whole equation: at 50 percent instead of 64, the residual climbs well past what any lawyer's normal read-through would reliably catch. Watch that number next, not the raw generation rate.
Naming the assumption that's both uncertain and consequential, not just the biggest number in the equation, is what a good estimator does that a bad one skips.
Two of the three flaw types sit exactly where a good estimator stops worrying: dangerous, but solved by a lookup. The third sits where the real work still is.
Three things worth naming directly, since this is where the real judgment sits. The AI-specific failure mode here is a mischaracterized holding: a real case, cited for a point it never actually decided, which is a specific and severe way for a legal tool to be wrong, not a generic bug. The guardrail is layered, a deterministic lookup for existence and staleness, plus a separate verification model for holding match, with every flagged citation and every high-value argument routed to a lawyer regardless of score. The alternative Falconbridge rejected was interviewing litigation partners for a target accuracy number instead of running the spike. It lost because nobody in that room had seen a real Holdfast mistake yet, so any number they gave would have been a guess wearing a requirement's clothes. And there's a real trade-off accepted on purpose: routing every flagged and every high-stakes citation through mandatory review means citations that could return in seconds now wait on a lawyer's actual time, slower and costing real billable hours, a cost Fairhurst and Merriwether accept because the alternative is a fabricated citation reaching a judge's desk.
And if you want to be sure it really works, try it somewhere else
Same five letters, a paper mill instead of a law firm. This time the unknown isn't a mischaracterized holding. It's whether a rare, quiet signal even exists inside years of sensor noise.
Different building, same shape of question. A mischaracterized holding became a bearing about to fail, and the method transferred without changing.
Faultline, built by Wrycroft Industrial, reads years of vibration, temperature, and pressure data off the slurry pumps at Winslowe Paper Co. and tries to flag a bearing failure before it happens. Nairne Oleander, the plant's reliability engineering lead, doesn't know if the failure signature is even detectable inside all that noise until someone actually builds and backtests a model. Nobody can spec that in a meeting.
Nairne runs the same five letters. Break it down: catching a failure only matters if the crew gets at least 48 hours of warning to schedule a repair without an emergency shutdown, so the spike needs to test against real, confirmed failures, not just "did the model notice something." Own the numbers: 40 confirmed pump failures over two years, full sensor history for each, backtested over four weeks by one engineer against the plant's own maintenance logs. Use a range: the model catches somewhere between 26 and 33 of the 40 failures with at least 48 hours warning, 65 to 82.5 percent, while throwing somewhere between 3 and 9 false alarms a month across the fleet of 600 pumps, a range because tightening the threshold to catch more failures always means more false alarms too. Nail the sanity check: an unplanned failure costs about $40,000 in downtime and damage, a false alarm costs about $800 in an unnecessary inspection, so even at the low end of detection and the high end of false alarms, the math still lands strongly in Faultline's favor. Direction: the swing assumption isn't the failure count, it's the false-alarm rate at whatever threshold catches enough real failures, because a crew that gets paged too often stops trusting every alert, the same way Jerric might have stopped reading a paragraph if the rhythm had kept holding.
The decision Wrycroft would take back
The first backtest just asked "did the model flag it before failure, yes or no," at one fixed threshold. That hid the real trade-off, more warning against more false alarms, inside a single number. Nairne would test multiple thresholds from day one instead of picking one and reporting a single hit rate.
Same rank, different lever, mapped onto BOUND: the breakdown is the same shape, four numbers hiding behind one word, unknown. The owned numbers are different but reasoned the same way, a real sample, a stated budget, a stated definition of "caught." The range is the same kind of honesty, an estimate from a backtest, not a census. And the swing number plays the same role in both stories: not the flashiest figure, but the one number that decides whether the team trusts the flagging enough to actually use it.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: how much of the outcome the model needs to get right, what it actually does on a real labeled sample, what a verification step catches, and what's left over compared to what the real thing can survive.
Cost: no budget for a full labeled sample this quarter. Run a smaller spike, 40 to 50 citations instead of 150, say so plainly, and treat the range as wider until a bigger one is affordable.
The model got better, for real: say the next version's raw mischaracterization rate drops to 3 percent instead of 9.3. The verification layer doesn't get removed. A lower raw rate isn't the same as zero, and the cost of one that slips through unflagged hasn't gotten any smaller.
Where people run it wrong.
They report the overall clean rate as the finding and stop, without ever isolating the one flaw type that actually needs a person.
They treat a small spike's range as a flaw in the method instead of the honest limit of what a three-week sample can tell you.
They test on the easy, obvious errors and call the feature validated, without checking whether the hard, judgment-only failures were even in the sample.
How to use it live. Ask yourself out loud, before naming any number: "which part of this is a lookup, and which part actually needs a person to decide?" That question alone buys real thinking time, and it's usually exactly what the interviewer is listening for.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
BOUND: turn a feasibility question nobody can answer in a meeting into one number, or a range, that decides whether to build, narrow, or kill it. Built for estimation and architecture questions, not a habit-flip story.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Vianne Underhill, who owns Holdfast's rollout at Falconbridge; Jerric Stanmore, the contract attorney who scored the spike's 150 citations by hand; and Zdenka Woolgar, the managing partner at Fairhurst and Merriwether who needs a real number before she'll trust it.
3 · THE BREAKDOWN
What three flaw types does the spike score separately, and why not one blended score?
Tap to flip
ANSWER
Fabricated (doesn't exist), mischaracterized (real case, wrong point), and overruled (real case, no longer good law). Blended into one score, 84.7 percent clean would have hidden that only mischaracterization actually needs human judgment to catch.
4 · THE SWING NUMBER
Which single number decides whether Falconbridge builds, narrows, or kills the generate-and-cite feature?
Tap to flip
ANSWER
The verification layer's catch rate on mischaracterized citations, measured at 64 to 86 percent. Existence and staleness are already solved by a lookup, so this is the only number still in real question.
5 · THE OLD DECISION
What decision would the team take back?
Tap to flip
ANSWER
Scoring the spike with one right-or-wrong column instead of splitting flaw types from day one. It was faster to set up, but it would have hidden the mischaracterization catch rate inside an average that sounded fine.
6 · THE NUMBER
Fill in the blank: of 150 citations, ___ were fabricated, ___ were mischaracterized, and verification catches only ___ to ___ of those.
Tap to flip
ANSWER
6 fabricated, 14 mischaracterized, and verification catches 9 to 12 of those 14. That means 2 to 5 citations out of 150 would reach a lawyer's screen looking clean when they weren't.
7 · THE REPLAY
Same three-week spike, flaw types split from day one, what changes?
Tap to flip
ANSWER
By day 12, nine days early, Vianne already had the number that mattered instead of waiting for all three weeks to produce one average. She brought Zdenka a real range, not a guess dressed as a passing grade.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs BOUND again on a different product. Which one, and what's the equivalent swing number?
Tap to flip
ANSWER
Faultline, Wrycroft Industrial's pump-failure predictor. Its swing number is the false-alarm rate at whatever threshold catches enough real failures, since pushing detection too high without limit trains crews to ignore every alert.
Check yourself Score: 0 / 0
Multiple choice
1. Why does the spike score citations into fabricated, mischaracterized, and overruled separately, instead of one pass or fail column?
A. Because attorneys refuse to work with a single column.
B. Because fabricated and overruled citations are basically solved by a database lookup, and blending them in hides that mischaracterization is the one flaw type that actually needs judgment and decides the answer.
C. Because Holdfast's own confidence score can only be split three ways.
D. Because the firm's malpractice insurer requires exactly three categories by name.
Show hint
Look at the B step, break it down, in the BOUND recap.
Show answer
B. Existence and staleness checks are close to deterministic. Mischaracterization is the only flaw type that needs a person's or a model's real judgment, and that's the one the whole discovery question turns on.
True or false
2. True or false: Holdfast's overall 84.7 percent clean rate is the number that should decide whether Falconbridge ships the generate-and-cite feature.
True
False
Show hint
Check the D step, direction, in the BOUND recap.
Show answer
False. The swing number is the verification layer's catch rate on mischaracterized citations, 64 to 86 percent. Existence and staleness are already solved by a lookup, so the overall clean rate hides the one number that actually decides the outcome.
Fill in the blank
3. Of the 150 citations in the spike, ___ were fabricated, ___ were mischaracterized, and ___ were overruled but not flagged. The verification layer's catch rate on the mischaracterized ones fell somewhere between ___ and ___ percent.
Show hint
Check the chart in Let's learn, and the U step in the BOUND recap.
Show answer
6 fabricated, 14 mischaracterized, 3 overruled, 64 to 86 percent. That range, pushed through the math, is what produces the 2-to-5 residual the direct answer is built on.
Short answer, name the rejected alternative
4. What alternative did Falconbridge consider instead of running a three-week technical spike, and why did it lose?
Show hint
Look at the "three things worth naming directly" paragraph near the end of the BOUND recap.
Show answer
Model answer: Interviewing litigation partners and writing a spec that named a target accuracy number, like 98 percent, without building anything first. It lost because nobody in that room, including the partners, had seen what Holdfast actually got wrong yet. A number produced from opinion is a guess wearing a requirement's clothes.
Short answer, apply it yourself
5. Think of an AI feature you've used or heard pitched where nobody could really know if it would work until someone built it. What's the one number a three-week spike on that feature would need to produce first?
Show hint
Ask what the B step asks: what does the discovery question actually need broken into parts, and which part is genuinely unknown.
Show answer
Model answer: An AI tool that drafts insurance claim denial letters, pitched as "explains every denial in plain language." Nobody knows until it's built whether the model can state the actual policy reason correctly, not just write something that sounds like a reason. A spike would need one number: on a labeled sample of real denials, what share of the AI's stated reasons match the real policy clause a human adjuster used.
Short answer, work the number
6. If a larger validation later showed the verification layer's real mischaracterization catch rate was 55 percent instead of the spike's measured 64 to 86 percent range, about how many of the 14 mischaracterized citations in a 150-sample would slip through completely unflagged?
Show hint
Residual equals 14 times one minus the catch rate.
Show answer
About 6, 14 times 0.45. That's worse than the top of the original 2-to-5 range, which is exactly the kind of result that would push the decision from proceed toward pivot or kill.
Before you close the answer
Why this works
Tests whether you'll actually build the narrow, dangerous slice first and measure it, instead of writing a spec with a number nobody can defend. Most candidates either skip straight to a full build plan or hand back an opinion dressed as a requirement.
Follow-up traps
"Why not just require attorney sign-off on every citation forever, and skip the whole verification-layer question?" Response: that's a real fallback, but it gives up almost all of the time saving Holdfast exists to deliver. It's the answer only if the catch rate can't clear roughly 50 percent even after more work, not the default.
"Isn't 150 citations too small a sample to build a real decision on?" Response: it's small on purpose. Three weeks and half of an attorney's time is what the team could actually spend before wasting runway or shipping blind. The 64-to-86 range is the honest admission that a bigger sample could move it, which is exactly why the plan calls for confirming it on more data before a full launch, not a reason the spike itself was wrong to run at this size.
If pressed
The verification layer's semantic check isn't the same model that generates the citation. It's a separate, smaller model trained specifically to compare a cited proposition against the actual holding text, scored independently, so a citation can never grade its own homework.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.