InterviewIntermediateModel Fluency & the AI PM Role / Managing stakeholder expectations and AI hype / #8

Explain to a non-technical executive why the last ten percent of quality costs more than the first ninety.

BOUND · pricing the last three points of accuracy for a non-technical executive, at Scrivano, an AI tool that transcribes legal depositions

Scrivano listens to a recorded legal deposition and produces a speaker-tagged draft transcript for a court reporter to check and certify. Morwenna Pettigrew is the product manager who owns its transcription model. Hadleigh Coldicutt is Scrivano's chief revenue officer, and he has never once opened a training log. Threadgold and Sconce, a litigation firm, runs its whole deposition practice through Scrivano, and Ambrogio Featherwick, the firm's director of litigation support, is the one whose court reporters caught a line that almost went out under the wrong witness's name.

The direct answer
Quality does not climb in a straight line. It climbs fast while the model is soaking up common, familiar patterns, then it slows hard once only rare, scattered failures are left, because each of those needs its own separate fix, not more of the same training pass. Show the executive the real cost per point, not just the percent, and let that number decide whether chasing the last few points is worth the months it actually takes.
Do this, in order
  1. Show the curve, not just the percent, before you promise a number.Why: a straight line and a bending one look the same on a dashboard until somebody has to defend the number out loud.
  2. Own the real cost per point, in dollars and weeks, not just in percent gained.Why: 2,700 dollars a point sounds nothing like 103,300 dollars a point, and only one of those is what's actually left.
  3. Break the last stretch into its real, separate causes before anyone quotes a date.Why: six unrelated problems don't get fixed by one more push of the same training.
  4. Give a range tied to how many causes are left, not one flat number for every case.Why: a routine deposition and a five-expert malpractice case are not sitting on the same curve.
  5. Check any straight-line promise against the real curve before it leaves the building.Why: this is the one check that would have stopped Hadleigh's call before it started.
  6. Decide out loud whether the last points are worth chasing, or the current bar plus a human check is smarter.Why: a real decision beats a wish that quality, speed, and cost could all keep improving for free.

How to answer this, stage by stage

Nobody's grading whether you can say "word accuracy" with a straight face. They're grading whether you can turn that number into something a revenue-focused executive can actually act on.

01
Scope it to one product and one promise
Say it like this
"Let's make this real. Scrivano transcribes legal depositions, and our chief revenue officer just told a client we'd hit 99 percent by next quarter because last quarter climbed fast. That's the promise I'm going to check, not accuracy curves in general."
Why this works
Pins a broad, abstract question to one real number instead of letting it drift into a lecture on model quality.
02
Name the method before touching a number
Say it like this
"I'd size this with BOUND. Break down what's actually left to fix. Own the real numbers behind the last push. Give a range instead of a lucky guess. Check the range against something real. Then say which single fact would move it most."
Why this works
Two seconds of structure tells a non-technical listener a real estimate is coming, not a hunch dressed up as one.
03
Reframe what the last few points actually are
Say it like this
"This was never really about the model getting slower. It's about what's left changing shape on us. The first ninety percent is one big, common problem. The last few points are six small, unrelated ones."
Why this works
This is the sentence that keeps the whole answer from sounding like an excuse about the model underperforming.
04
Break down what's actually inside the last stretch, the B step
Say it like this
"Crosstalk when two lawyers object at once. Rare or non-native accents. Expert jargon from a witness's own field. Two people talking over each other. Bad mic pickup. A handful of one-off cases nobody's seen twice. None of those share a fix."
Why this works
A non-technical listener can hold six named things in his head. He can't act on "the model is at 95 percent."
05
Own the real numbers, the O step
Say it like this
"Going from 80 to 95 percent took ten weeks and one 40,000 dollar batch of ordinary deposition audio. Going from 95 to 98 took twenty weeks and 310,000 dollars, because we were paying for six separate fixes instead of one big one. That's about 38 times more, per point."
Why this works
A number with its source survives a follow-up question. "It's harder now" does not.
06
Give a range, and name what makes it steeper, the U step
Say it like this
"A routine, single-witness contract deposition sits on the flat part of this curve, maybe one or two leftover causes. A five-expert malpractice case sits on the steep part, all six at once. The number of separate causes left is what decides which one you're on, not how hard the team is working."
Why this works
Turns "it depends" into something the executive can check against his own client's actual case load.
07
Run the sanity check before the number leaves the room, the N step
Say it like this
"If ten weeks bought fifteen points, it's tempting to think another ten weeks buys the next fifteen. It bought three, not fifteen, at eight times the cost. The first ten weeks fixed one common problem. This twenty weeks was fixing six rare ones."
Why this works
This is the exact line that would have stopped the promise before it went out on a client call.
08
Name the direction, then close on the real decision, the D step
Say it like this
"The single fact that moves this most is how many separate causes are left, not how good the model is or how hard the team pushes. So here's the actual choice. Spend five months and 310,000 dollars chasing three more automated points, or hold the model at 98 and let a person catch the rare cases it still misses. Either one is defensible. Promising 99 by next quarter without saying which one we picked is not."
Why this works
Ends on a decision the executive can actually make, not a feeling that the team is trying hard.

Let's learn

Hand sketched left to right flow diagram titled How a deposition becomes a certified transcript. Five connected boxes reading Deposition recorded, Scrivano transcribes and tags speakers, Low-confidence lines get flagged, this box outlined and emphasized in rust orange, Court reporter checks the flagged lines, Transcript certified and filed.
Five boxes. The third one, flagging what the model is unsure about, is the whole reason the last five points are survivable at all.

Scrivano is a button a court reporter presses when a deposition wraps for the day. It listens to the recording, writes out every word, and tags who said it, so the draft is mostly done before anyone opens a laptop.

Before Scrivano, turning a four hour deposition into a rough draft took a court reporter about nine hours of listening and typing. With Scrivano running at its first real number, 80 percent word accuracy, that draft is ready in about fifteen minutes, though the reporter still spends close to five hours fixing what it got wrong. Ten weeks later, once the model reached 95 percent, that clean up time dropped to about forty minutes, mostly spent checking the handful of lines Scrivano itself flags as unsure.

Knowledge spark: what does word accuracy actually count? It's the share of words in the transcript that match what was really said, checked against a human-verified copy. A model can sit at 95 percent overall and still be badly wrong on one specific kind of sentence, because the miss isn't spread evenly across the page.

Here's the turn. The five points Scrivano still misses at 95 percent are not spread evenly across an ordinary transcript. They cluster in the exact moments a legal record gets argued about hardest: two lawyers talking over each other during an objection, a witness testifying in a second language, a specialist using words a general model has barely seen before.

Hand sketched numbered icon list titled What is actually left in Scrivano's final five points. Six rows: one, two people icon in rust orange, Crosstalk when two lawyers object at once. Two, a question mark box icon in amber, Rare or non-native witness accents. Three, a document icon in amber, Expert witness technical jargon. Four, a scale icon in rust orange, Overlapping speakers, wrong line to wrong name. Five, a gauge icon in warm grey, Microphone bleed-through on the recording. Six, a funnel icon in warm grey, One-off causes nobody has seen twice.
Six separate causes, not one blurry five percent. Each row needs its own fix, its own data, its own timeline.
We didn't lose five points of accuracy. We lost the five points sitting inside the four seconds a courtroom actually fights over.

What it costs at its worst: during a deposition of a structural engineer testifying as an expert witness, two attorneys spoke over each other during an objection. For about four seconds, Scrivano's draft briefly put a damaging line in the wrong person's mouth. The line scored low confidence, so it got flagged, and a senior court reporter caught it before certification, exactly the way the system is meant to work. Nothing went out wrong. It was close enough that the client called the same week Hadleigh was writing his renewal pitch.

The choice I would take back Early on, with one client and one kind of case, Scrivano's team built a single accuracy dashboard, one blended line across every deposition type and every client. That was a fine, honest choice with one client. It became a liability the moment somebody read a promise off of an average. I'd take it back and split that dashboard by case complexity from day one, so sales can't accidentally sell off a number that was never really one number.

What I would leave alone: a routine, single-witness commercial contract deposition, clear audio, plain vocabulary, sits on a flat part of this same curve. Pushing that kind of case from 95 to 98 percent is genuinely cheap, because almost everything left in it is one or two causes, not six. That's worth chasing hard. The expensive climb only shows up where the causes pile up.

The lesson: a confident number on a dashboard and an honest one look exactly the same, right up until somebody has to defend it out loud, on a call, to a client deciding whether to renew.

Now here is the same thing as a story

The short version above is what you'd actually say out loud in an interview. Read this one for the Thursday the promise almost went out wrong twice in the same week.

Scrivano's accuracy dashboard lives on a shared screen in the sales bullpen, and by nine most mornings, Hadleigh Coldicutt has already glanced at it twice. He's not an engineer, and he's never pretended to be one. What he's good at is closing a room, and for the past year the dashboard had been doing half his job for him.

Morwenna Pettigrew has run Scrivano's transcription model since before it had a single paying client. She reads an eval report the way a proofreader reads a page, catching a number drifting off before anyone else has noticed it moved.

January was launch month, 80 percent word accuracy, rough but real. By mid-March, ten weeks in, the line had climbed to 95, and Hadleigh had started opening every sales call with a screenshot of that climb. New firms signed partly because the line kept going up.

The habit thinned out in three small steps. In February, on a call with an existing client, he said "it keeps getting better every month," true, harmless, and he'd checked the rough shape of it with Morwenna first. In April, pitching a new firm, he said "give us a quarter and we'll be near-perfect," an extrapolation now, not a fact, and he didn't check with anyone before he said it, though it happened to still sound roughly right. By June, prepping the renewal deck for Threadgold and Sconce, he typed "99 percent by Q4" straight into a slide, a specific number and a specific date, and he never asked Morwenna a single question about it.

Then, on a Thursday, Ambrogio Featherwick called Scrivano's account team, not angry, just unsettled. One of his senior court reporters had caught a line during a routine review: four seconds of crosstalk during an objection, a damaging statement briefly attributed to the wrong person. The system had flagged it low confidence, exactly as designed, and nothing had gone out wrong. But Ambrogio couldn't stop thinking about the version of that afternoon where nobody happened to check that particular line.

Hand sketched three panel comparison titled The Thursday the renewal call almost went wrong. Left panel, a gauge icon in slate blue, labeled Hadleigh preps the call, caption reads the dashboard, sees one straight climb. Middle panel, a question mark box icon in rust orange, labeled Ambrogio's near miss lands, caption a crosstalk line almost named the wrong witness. Right panel, a document icon in forest green, labeled Morwenna steps in, caption brings the real cost per point before the call starts.
Three things landing on the same Thursday. Only one of them was actually a coincidence.

Morwenna heard about the near miss at 1pm. Hadleigh's call with Ambrogio's firm, the renewal call, the one with "99 percent by Q4" already sitting in the deck, was set for 5pm the same day.

She spent the afternoon building the one chart she should have built months earlier: what it actually cost, in dollars and weeks, to move each point past 95. By 4:45pm she was in front of Hadleigh with it.

Hand sketched two panel comparison titled The line Hadleigh drew versus the curve Scrivano actually walks. Left panel, a gauge icon in slate blue, labeled Hadleigh's line, caption reads same pace every ten weeks, straight to 100. Right panel, a scale icon in rust orange, labeled Scrivano's real curve, caption reads fast to 95, then it bends hard and slows.
Both lines start in the same place. Only one of them describes what actually happened next.
Hadleigh wasn't wrong that the line had been climbing for months. He was wrong about what was left to climb.

He remembered the January meeting where the single blended dashboard got decided. "Nobody in that room was being lazy," Morwenna told him. "There was one client. There was one kind of case. A blended line was the honest line, then." It just stopped being honest the moment a second, harder client came along and nobody split the number back apart.

At 5pm, Hadleigh got on the call. He didn't say 99 percent by Q4. He told Ambrogio the real curve, the six causes, the near miss, and the actual cost of chasing the last points automatically. Then he offered something better than a guess: 99 percent on Threadgold and Sconce's routine single-witness matters within the quarter, since that segment was cheap to push, and a slower, staged target for the multi-expert cases, backed by the same human review that had just caught the crosstalk line.

The renewal closed anyway, two years, at 840,000 dollars a year. And the following week, the segmented dashboard Morwenna should have built in January finally got built, this time before anyone needed it to survive a phone call.

What she'd tell her January self: she built that dashboard for engineers to trust. She never once asked what happens the day somebody in sales reads a promise straight off of it.

BOUND, for pricing the three points Hadleigh thought were free

Not a way to make a bad number sound careful. BOUND turns "it's getting better" into a cost per point Hadleigh, or anyone else at Scrivano, could actually stand behind on a client call.

BBreak it down. What is actually left, and why doesn't one fix cover it?
"Quality" isn't one number. It's the sum of everything the model still gets wrong, and past 95 percent, that sum splits into six unrelated pieces: crosstalk during objections, rare or non-native accents, expert jargon, overlapping speaker mix-ups, mic bleed-through, and a long tail of one-off cases. The first 90 percent came from one shared thing, common, well-recorded, plain-English testimony that shows up thousands of times in training audio. The rest doesn't share a cause, so it can't share a fix.
Skip this split and "get to 99 percent" sounds like a single push, when it's really six separate small projects wearing one number's clothes.
OOwn the numbers. Where does each one actually come from?
Ten weeks and one 40,000 dollar batch of broad, ordinary deposition audio took Scrivano from 80 to 95 percent, 15 points, about 2,700 dollars a point. Twenty weeks and 310,000 dollars took it from 95 to 98, only 3 points, about 103,300 dollars a point, because that money bought six narrow, targeted fixes instead of one broad pass.
A number only counts as owned if you can say exactly where it came from when someone pushes on it. "It's harder now" is a feeling wearing a number's clothes.
Weeks invested versus word accuracy: the promise against the real curve
75% 85% 95% wk0 wk10 wk20 wk30 wk40 wk10, 95%, broad tuning done Hadleigh expected 100% by wk13 wk30, 98%, targeted fixes done
Real accuracy curveThe straight-line promise
Both lines leave week 10 in the same place. By week 20 the real curve is at 97.2 percent and the promised one is already sitting at 100.
UUse a range, not one number.
A routine, single-witness contract deposition, clear audio, plain vocabulary, sits on the flat part of this curve, one or two leftover causes at most. A multi-expert case, a deposition with several specialists and heavy crosstalk, sits on the steep part, all six causes at once. What makes the curve steeper isn't the case, it's how many separate, unrelated causes are still sitting in what's left.
A single number here repeats Hadleigh's exact mistake: sounding certain about something that genuinely depends on which kind of case is being asked about.
Hand sketched horizontal timeline titled How long the last three points actually take. Four marked points along the line: Best case, 14 weeks, one or two narrow causes. Typical, 20 weeks, six separate causes, 310 thousand dollars. Worst case, this point emphasized in rust orange, 30 plus weeks, still short of 99 percent. Sanity check, the whole 80 to 95 push took only 10 weeks.
Four points on one line. The gap between best case and worst case is entirely about how many causes are actually left, not how hard anyone works.
NNail the sanity check. Does the number survive contact with the last quarter's own numbers?
If ten weeks bought fifteen points, the tempting math says another ten weeks buys the next fifteen, or close to it. It didn't. Twenty weeks, twice as long, bought three points, one fifth as many, for nearly eight times the money. The first ten weeks fixed one big, common problem. The next twenty were spent fixing six small, rare ones, one at a time.
This is the exact check Hadleigh's deck skipped, and it's the one line that would have stopped the promise before it reached a client.
Cost per accuracy point, broad tuning versus targeted fixes
about $2,700 / point Broad tuning common patterns, 15 points, $40,000 about $103,300 / point Targeted fixes six rare causes, 3 points, $310,000
Broad tuning, weeks 0 to 10Targeted fixes, weeks 10 to 30
Same team, same product, same ten weeks either way. The only thing that changed was how many unrelated causes each dollar had to chase.
DDirection. Which assumption moves this most, and what's the actual decision?
Not how good the model is, and not how hard the team pushes. The single biggest swing factor is how many separate, unrelated causes are still sitting in whatever's left. So the real decision isn't a date. It's a choice: spend five months and 310,000 dollars chasing three more automated points, or hold the model at a strong, provable 98 percent and pay a person to catch the rare, expensive cases it still misses, for as long as Scrivano runs.
Naming the fact that actually swings the estimate, instead of the biggest number in it, is what separates a real estimator from a confident guesser.
Hand sketched two panel metaphor scene titled Why the last stretch of any quality job costs more. Left panel, a box icon in forest green, labeled Framing the shell, caption covers most of the square footage, weeks not months. Right panel, a scale icon in rust orange, labeled Fitting the crown molding, caption one crooked doorway at a time, a specialist, months.
The analogy that actually lands with someone who has never opened a training log: renovating an old building costs the same way quality does.

One alternative considered and rejected: hire enough human proofreaders to check one hundred percent of every transcript by hand, instead of continuing to fund model fixes. It lost, because it scales the wrong way. More depositions would mean more hours, forever, instead of a fix that keeps paying off once it's built. The AI-specific risk sitting under all of this is quiet: a wrongly attributed line inside four seconds of crosstalk reads exactly as clean and confident as a correct one, with nothing on the page to warn a tired reviewer. The guardrail is the per-segment confidence score, tied to a mandatory human check on anything flagged low-confidence or overlapping, before certification, the exact thing that caught the near miss. And the trade-off is real, not free: five months and 310,000 dollars to chase three more automated points, against holding the model at 98 percent and paying a person to catch the rare cases for good.

And if you want to be sure it really works, try it somewhere else

Same five letters, a crop-disease photo checker instead of a deposition tool, and this time the executive's real question isn't whether the curve bends, it's whether bending it is worth the money at all.

Loamwright checks a phone photo of a crop leaf and flags the ones likely to be diseased, so a farmer doesn't wait for a spreading blight to become obvious to the eye. Evrart Stonemere is Loamwright's product manager. Longbarrow Growers Collective, a farm co-op, runs it across its member farms, and Guntram Ashencroft, the co-op's field operations director, wants near-certain detection before trusting it across every farm in the collective.

Run BOUND on it. Break it down: after broad tuning on common blights and rust in good daylight, what's left splits into blurry or motion-blurred photos, a rare strain nobody's photographed much, backlit glare hiding a leaf's true color, and overlapping leaves hiding the actual spot. Own the numbers: eight weeks and 25,000 dollars took accuracy from 85 to 96 percent, 11 points. Sixteen weeks and 180,000 dollars took it from 96 to 98.5, only 2.5 points, over seven times the cost per point.

Hand sketched labeled parts diagram titled What is left in Loamwright's last two points. A gauge icon labeled The remaining misses sits in the center, with four labeled callouts around it: Blurry, motion blurred photos. A rare disease strain. Backlit glare on the leaf. Overlapping leaves hide the spot.
Different modality, same shape of problem. A photo model's last few points splinter into named causes exactly the way an audio model's do.
Where Loamwright's answer genuinely differs Morwenna's gap was a promise made without checking the model's owner first. Evrart's team already knew to check before promising anything, agricultural AI teams learn that fast. Their real question was different: whether the 180,000 dollars was worth spending at all, since most of Longbarrow's member farms only ever photograph the two most common blights, right where this curve barely bends.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: split what's left into its real, separate causes, own a real cost for the common part, and turn the unknowable rare part into a checked range, never a guessed one.
Cost: no budget to hold a promise past one month. Shrink the range honestly instead of hiding it, and say plainly what "good enough" has to mean by then.
The model got better, for real: say Loamwright's overall accuracy jumps to 99 percent right as the promise is being made. Still check by segment, since a strong headline number can hide a badly-performing slice just as easily as a weak one, maybe more easily, since nobody goes looking underneath good news.

Where people run it wrong.
They let "it's getting better" stand as the whole plan, instead of naming what's actually left.
They read one dashboard average and call the model even everywhere, instead of checking case by case, or farm by farm.
They treat a fast early climb as proof the next climb will move at the same pace.

How to use it live. Before answering, ask yourself one plain question out loud: "is what's left one shared problem, or several separate ones." Whichever it is, that answer decides whether the next stretch is cheap or expensive.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What method fits explaining why the last few points of quality cost so much more than the first ones?
Tap to flip
ANSWER
BOUND: break it down, own the numbers, use a range, nail the sanity check, name the direction. Built for turning a fuzzy timeline promise into numbers that survive a follow-up question.
2 · THE CAST
Who is this answer about?
Tap to flip
ANSWER
Morwenna Pettigrew, Scrivano's transcription PM. Hadleigh Coldicutt, Scrivano's non-technical chief revenue officer. Ambrogio Featherwick, director of litigation support at Threadgold and Sconce, the client firm.
3 · THE BREAKDOWN
What's actually left in Scrivano's final few points, and why doesn't one fix cover it?
Tap to flip
ANSWER
Six separate causes: crosstalk during objections, rare or non-native accents, expert jargon, overlapping speaker mix-ups, mic bleed-through, and one-off cases. Each needs its own data and its own fix, not one more broad training pass.
4 · THE OWNED NUMBER
Fill in the blank: ___ weeks and $___ took Scrivano from 80 to 95 percent. ___ weeks and $___ took it from 95 to 98 percent.
Tap to flip
ANSWER
10 weeks and $40,000. Then 20 weeks and $310,000, about 38 times more per point.
5 · THE RANGE
What decides whether a deposition sits on the steep part of this curve or the flat part?
Tap to flip
ANSWER
How many separate causes are left in its remaining errors. A routine single-witness case might have one or two. A multi-expert case has all six. More unrelated causes means a steeper curve.
6 · THE SANITY CHECK
Why was Hadleigh's straight-line math wrong?
Tap to flip
ANSWER
He assumed the next ten weeks would buy the same kind of progress as the first ten. The first ten weeks fixed one big, common problem. The next twenty were fixing six small, rare ones, so they bought three points, not fifteen, at eight times the cost.
7 · THE DIRECTION
Which single fact swings this estimate the most, and what's the real decision?
Tap to flip
ANSWER
How many separate causes are left, not how hard the team works. The real choice: spend five months and $310,000 chasing three more automated points, or hold the model at a strong number and pay a person to catch the rare cases it still misses.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs BOUND again on a different product. Which one, and what's the same shape?
Tap to flip
ANSWER
Loamwright, a crop-disease photo checker for Longbarrow Growers Collective. Same shape: a cheap broad push to 96 percent, then a slow, expensive push past it because what's left splits into blur, a rare strain, glare, and hidden leaves.

Check yourself Score: 0 / 0

Fill in the blank
1. Scrivano's word accuracy went from 80 to 95 percent in ___ weeks, for $___. Getting from 95 to 98 percent took ___ weeks, for $___.
Show hint
Check the O step and the cost-per-point bar chart's two totals.
Show answer
10 weeks, $40,000. Then 20 weeks, $310,000. The second push took twice as long and nearly eight times the money, for a fifth as many points.
Multiple choice
2. Why did the push from 95 to 98 percent cost about 38 times more per point than the push from 80 to 95?
  • A. The engineering team got slower and less skilled over time.
  • B. The remaining errors were spread across six separate, mostly unrelated causes that each needed their own fix, instead of one shared cause.
  • C. Threadgold and Sconce demanded a stricter contract.
  • D. Running Scrivano got more expensive as more depositions were processed.
Show hint
Check the B step and the icon list of six causes in Let's learn.
Show answer
B. One broad pass fixed one common problem. Six narrow, unrelated causes each needed their own separate fix, which is what made the second push so much more expensive.
True or false
3. True or false: once a model clears 95 percent accuracy, the fastest way to reach 99 percent is to run the exact same kind of broad training pass that got it from 80 to 95.
  • True
  • False
Show hint
Check the N step and the line chart's real curve against the dashed promise.
Show answer
False. That broad pass fixed the one big, shared problem. What's left past 95 percent is six small, separate, rare problems, and a broad pass barely touches any single one of them.
Short answer, apply it yourself
4. Think of something in your own work or life that got easy fast and then got hard slowly, learning an instrument, fixing bugs in old code, cleaning out a garage. What made the last stretch different from the first?
Show hint
Look for whether the last stretch was one shared cause or several unrelated small ones.
Show answer
Model answer: Learning to cook well. The first stretch, basic technique, timing, seasoning, comes fast from repeating the same handful of dishes. The last stretch is a different kind of hard, dozens of specific dishes each with their own one-off trick, and no single lesson covers them all.
Short answer, work the number
5. If Threadgold and Sconce had insisted on 99 percent within one more quarter no matter the cost, what would Morwenna's own numbers suggest about how long and how much that could really take?
Show hint
Look at the cost per point for the 95 to 98 push, and think about where the next point sits on that same curve.
Show answer
Model answer: Going from 95 to 98, three points, already took 20 weeks and $310,000. The next point sits on an even steeper part of the same curve, so a honest estimate is well past one more quarter, likely six months or more and several hundred thousand dollars, with real risk that some of what's left, like mic bleed-through, can't be fully fixed by the model at all.
Before you close the answer
Why this works
Tests whether a candidate can turn a true technical fact, that the tail of a quality curve is made of unrelated rare cases, into a cost and timeline argument a revenue-focused executive can actually act on, instead of hiding behind the word "accuracy."
Follow-up traps
"Can't you just throw a bigger model or more compute at it?" Response: a bigger model mostly helps with the common ninety percent, since that's a matter of pattern volume. It doesn't fix six unrelated, rare causes with no shared pattern between them, that needs targeted data for each one, not more scale on the same data.

"So should Scrivano ever promise 99 percent?" Response: yes, but as a number tied to a specific case type, like 99 percent on routine single-witness depositions where the curve is flat, not one blanket number across every kind of case a firm runs.
If pressed
Scrivano's confidence flag isn't a separate model. It's read straight off the transcription model's own token probabilities, and anything under 90 percent confidence, or any stretch where two speakers' audio overlaps for more than half a second, gets queued for a mandatory human check before certification, which is exactly what caught the near miss in the first place.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more