How do you value a feature that improves quality rather than throughput?
Cuewright is the AI that writes captions for every show Kestrelbourne Media streams, live and on demand. Silvana Colquist owns what Cuewright is actually worth. Oisin Thistlewaite, who runs Content Operations, reports one number for it every quarter: the share of episodes captioned inside four hours of upload. Then a new hire on the Accessibility Insights team, Ysanne Northspan, asked a question neither dashboard had ever been built to answer.
- Value the quality gain against complaints and churn avoided in the caption-dependent segment, never against the throughput SLA.Why: a word-error-rate improvement literally cannot move a four-hour turnaround number, so grading it against that metric always reads as "no ROI."
- Isolate the caption-specific complaint-to-churn gap before naming any dollar figure.Why: blending it with general accessibility churn is exactly the mistake that gets a claim corrected in front of finance.
- Carry the number as a range, and say "modeled" out loud.Why: it's built from a churn-gap correlation, not a subtraction from a real invoice, and one clean figure claims certainty nobody has.
- Treat the confidence threshold as a lever with a real price, not a free knob for hitting the SLA.Why: tightening it to protect on-time turnaround is what let 160 extra tickets a month, and the churn behind them, go unpriced for two quarters.
- Set real kill criteria for the churn-gap proxy, and actually check it.Why: a matched-cohort test that shows the gap collapsing to noise should move the case to a different proxy, not get ignored because the current story is convenient.
- Don't undersell the parts of Cuewright that already show up on an invoice or a timesheet.Why: real, checkable savings deserve credit too; downplaying them to look cautious just starves the case for the whole product.
How to answer this, stage by stage
Nobody is grading whether you know the term word-error rate. They're grading whether you can price a real win that a throughput dashboard structurally cannot register, and defend that price when someone tries to force it back onto the dashboard anyway.
Let's learn
A caption only has two settings for the person reading it: right, or wrong. There's no setting in between that still lets you follow the scene.
Cuewright is Kestrelbourne Media's captioning engine. It listens to a show's audio and writes the caption line, timed to the second, for everything Kestrelbourne streams, live or on demand.
Before Cuewright, an outside vendor captioned every new episode by hand. It worked, but captions landed three to four business days after an episode went up, at $4.20 a finished minute. Kestrelbourne kept a small standing team just to chase that vendor before big international premieres.
Cuewright drafts the same episode's captions in minutes now. Most of Kestrelbourne's 3,400 monthly hours of new programming publish captions within four hours of upload, the number Content Operations reports every quarter as Cuewright's headline win.
Here's the turn. The extra wrong words aren't really the story. The story is what a caption-dependent viewer does after hitting one wrong line in a scene they can't otherwise follow. Most of them don't write in to complain. A little over a third of the ones who do are gone within two months.
Kestrelbourne tags a subscriber as caption-dependent once they've kept always-on captions running for ninety straight days, not just switched them on for one loud restaurant. About 88,000 subscribers, a little over 4% of the platform, carry that tag. Of those who file an accuracy ticket, 34% cancel within sixty days. Subscribers in that same segment who never file one cancel at 6% over the same window. Multiply that 28-point gap by the 160 extra tickets a month and Kestrelbourne's own modeled subscriber value, $310, and the tightened threshold works out to about $166,656 a year, call it $135,000 to $205,000 once you allow for how uncertain that gap really is.
What it costs at its worst: two full quarters of churn nobody had priced, because the team celebrating 97% on time and the team logging caption complaints had never once opened each other's dashboards.
What I would leave alone: a filler word transcribed as "um" instead of "uh," or a comma landing one beat late. Those don't change what a viewer understands. Cuewright's threshold doesn't need to catch every one of those, and chasing them would just slow the review queue down for nothing.
The lesson: quality that never moves a throughput dashboard isn't quality with no value. It just needs a different ruler. Build that ruler on purpose, or someone finds the bill by accident, months later.
Now here is the same thing as a story
Read the short version above when you're in the room. Read this one when you want to feel why a new hire's first month, not a bad quarter, is what actually surfaced the bill.
Every Monday morning, before the rest of Content Operations has logged on, Silvana Colquist opens the same three dashboards. She's run Cuewright's product decisions for two years, since before Kestrelbourne trusted an AI model with a single caption. She knows which shows give it trouble before she opens the report: anything with invented names, anything with more than two people talking over each other.
The first year was good, plainly. Cuewright cut the vendor bill by more than Kestrelbourne had budgeted for, and the four-hour turnaround number climbed steadily, quarter over quarter, everyone's favorite kind of chart.
It thinned in three beats, and none of them looked careless. Beat one: Kestrelbourne signed two new content deals, and the volume of new programming grew by nearly half. Beat two: Content Ops, watching the four-hour SLA start to slip under the extra load, quietly moved Cuewright's confidence cut-off, so fewer borderline lines went to a person before publishing. Beat three: the SLA recovered, then beat its old best, 97% on time, and nobody went looking for what had paid for that.
The trigger wasn't a bad quarter. It was a new hire's first month. Ysanne Northspan joined the Accessibility Insights team in March, and one of her first tasks was a routine segment-health report, the kind nobody reads closely. She noticed two numbers sitting three tabs apart: caption-accuracy tickets from the caption-dependent segment, up from 180 to 340 a month, and that same segment's sixty-day churn, running far higher whenever a ticket was attached. Nobody on her team could tell her why. Nobody on Content Ops had ever been asked.
Silvana didn't have an answer either, not really. What she had was a dashboard that said Cuewright was having its best quarter ever, and a spreadsheet Ysanne had built in an afternoon that said something else entirely. So she did the thing nobody had done in six months: she put the two numbers next to each other, on purpose, and worked out what the gap actually cost.
I want to say the problem was Cuewright getting worse. It did get a little worse; word-error rate crept up once fewer lines got a second look. But that's not really the story. The four-hour dashboard never had a setting for "a caption-dependent viewer stopped trusting this." It only had a stopwatch.
The decision that opened the door traced back to a fifteen-minute stretch in a load-planning meeting, back when the two new content deals landed. Someone on Content Ops asked whether tightening the confidence cut-off would cost anything on the caption-quality side. The honest answer at the time was: nobody had a number to check it against. So the meeting moved the line, and moved on.
Run that meeting again, with one number added: recalibrating the cut-off back costs about $87,000 a year in extra review labor, and it's expected to hold on-time turnaround around 93%, not 97%. Against a modeled $135,000 to $205,000 a year in avoided churn, that's not a close call. Silvana runs the numbers, Content Ops moves the line back by nine the next morning, and by the following quarter's report, caption-accuracy tickets are already sliding back toward 180.
One design let "on time" mean whatever the confidence cut-off happened to be sitting at. The other lets it mean "on time, and checked against what it actually costs to move."
What Silvana would tell herself, back in that fifteen-minute meeting: the question was never "does this cost anything." It was "do we have a number to find out," and for six months, the honest answer was no.
PICK: choosing a ruler the dashboard doesn't have
Not a way to dress up "it's complicated" in four letters. PICK forces a real commitment about which ruler a quality gain actually earns, then makes you say, out loud, which kind of mistake actually costs someone money.
And if you want to be sure it really works, try it somewhere else
Same four letters, a mechanical fault instead of a wrong word, and this time the modeled dollar isn't a subscriber walking away. It's a technician driving out twice for one job.
Coilsense is Ferric Field Services' diagnostic AI. A technician photographs a unit and logs its sensor readings, and Coilsense scores whether the fault needs a truck roll now, later, or not at all. Ferric runs it across about 5,200 service calls a month.
Adalric Verhagen runs field operations at Ferric, and the metric his regional office watches is dispatch speed: the share of calls triaged within thirty minutes of the customer's photo landing. Six months ago, someone tightened Coilsense's auto-clear threshold to hit a new dispatch-speed target, letting more borderline "all clear" calls skip a technician's second look. Triage speed hit 91%. Missed faults, cases where a unit that actually needed work got auto-cleared, crept from about 2% to 5.5% of cleared calls, about 45 extra repeat truck rolls a month, roughly $9,400 in avoidable cost that never showed up on the dispatch-speed chart.
Same rank, different lever: the fix isn't a smarter model. It's pricing the false all-clear rate against the cost of a repeat truck roll, the same way Kestrelbourne priced caption tickets against churn, and treating the auto-clear threshold as a lever with a real cost on both sides.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: value the quality win against what it avoids downstream, never against a throughput number it can't move, and carry the figure as a range.
Cost: no budget this quarter for a matched-cohort study. Ship the cheap version first: a spreadsheet comparing last year's ticket or repeat-visit rate before and after the threshold moved, revisited next quarter with real data.
The model got better, for real: say Cuewright's word-error rate drops by half overnight. The valuation habit doesn't change. A lower error rate should shrink the avoided-churn range, not excuse skipping the recheck that confirms it actually did.
Where people run it wrong.
They let the throughput metric be the only ruler in the room, because it's the one that already has a dashboard.
They value the quality win with a single point number instead of a range, and get corrected the first time someone asks where it came from.
They watch the aggregate satisfaction score and miss that it's dominated by the majority who were never affected in the first place.
How to use it live. Ask the ruler question before naming a number: "Is this feature supposed to move the throughput metric, or does it need its own ruler because the group it helps is a small slice of the dashboard?" That question alone usually tells you whether you're about to defend a real number or force one that doesn't fit.
Three things worth stating directly, since this is where the real judgment sits. Silvana considered valuing the quality gain directly against the word-error-rate improvement, call it "WER dropped three points, so it's worth X," and rejected it, because a percentage point of word-error rate has no dollar value on its own. It only matters through what a viewer does after hitting a wrong caption, which is exactly why the churn-gap proxy had to exist instead. The AI-specific failure worth naming is that Cuewright's confidence score measures how sure the model is about the sounds it heard, not whether the word is the right word: a fluent, made-up-sounding character name can score just as confident as a real one transcribed correctly, so the threshold catches mumbled audio far better than it catches a wrong but fluent name. The guardrail is a separate golden-set eval, built from Kestrelbourne's most jargon-dense shows, hand-scored and tracked across every model version, independent of the per-line confidence routing, since that routing structurally can't catch this exact class of error. And the trade-off is real: recalibrating the threshold back costs about $87,000 a year in extra review labor and settles on-time turnaround around 93% instead of 97%, a real cost Kestrelbourne accepted on purpose once the churn number made the case.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Why not just require every caption to get a human check, and skip the confidence threshold entirely?" Response: reviewing every line would cost far more than the whole review budget already does, for a turnaround time no viewer, hearing or not, would accept; the threshold is the trade Kestrelbourne chose, recalibrated with a real price on both sides instead of protecting one side for free.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Measuring ROI and business impact
- #1 How do you build the ROI case for an AI feature before it ships?
- #2 What is the difference between time saved and value created?
- #3 Model the annual ROI of a support agent that deflects 30 percent of tickets.
- #4 How do you attribute a revenue change to an AI feature specifically?
- #5 Explain why time-saved metrics are frequently overstated.
- #6 Describe an experiment design that would isolate an AI feature's business impact.