CalculationAdvancedQuality, Cost & Token Economics / Cost modeling and unit economics / #14

What is the cost of your eval and monitoring infrastructure relative to serving?

ORDER · eval and serving cost

Reedline's eval and monitoring budget was never supposed to track serving dollar for dollar. It was supposed to shrink on purpose. The one time it shrank the wrong way, a guest's company name came back wrong in three sponsors' compliance reports before anyone caught the pattern.

The direct answer
Spend more on eval and monitoring than on serving while a product is small. Catching a wrong name before it reaches a sponsor's compliance report is the one decision you cannot undo once it ships, and early on that protection should cost several times what serving costs, for Reedline it started at eleven times. Let the ratio fall as volume grows, but only by making the eval harness cheaper to run, never by reviewing the same thin slice of episodes no matter who is in them. At full scale it should settle near a floor and stay there. For Reedline that floor sits around seventeen percent of serving cost, never zero.
Rank the budget, in order
  1. Spend more on eval and monitoring than on serving while the product is small, and never let that spend hit zero once it isn't.Why: a wrong name in a sponsor's compliance report costs a client relationship. A slower processing queue costs an apology email. Only one of those is hard to undo.
  2. Let the eval-to-serving ratio fall as volume grows, but only by making the eval harness cheaper to run, never by reviewing the same slice of episodes no matter who is in them.Why: the flat one-in-five human sample that worked at 240 episodes a month quietly under-caught the riskiest one in eight once volume hit 14,000.
  3. Build the golden set and an automated eval harness before a network's volume scales past pilot size.Why: you can't safely carry 90,000 episodes a month on review capacity sized for a 240-episode pilot.
  4. Route the shrinking review budget toward the riskiest slice, a new host's first weeks, instead of spreading it evenly across every episode.Why: an even review rate spends most of its budget re-checking episodes that were already accurate.
  5. Run a small routine audit that recuts the error rate by cohort, not just the blended average.Why: the blended number sat at a healthy looking 2.2 percent the whole quarter new-host episodes were already failing at 13.4 percent.
  6. Fund monitoring ahead of serving-capacity growth whenever the two compete for the same budget.Why: adding capacity back takes a purchase order. Winning back a sponsor's trust after a wrong name ships takes months, if it happens at all.

How to answer this, stage by stage

Nobody is grading whether you can name a percentage. They're grading whether you can say why that percentage should be different this year than it was last year, and different again next year.

1
Scope it to one concrete product before ranking anything in the abstract
Say it like this
"Let's ground this in one tool. Reedline is Briarcask Audio's AI product. It listens to a podcast episode and drafts a transcript plus show notes, so a producer doesn't have to write them by hand after every episode. Esben Brandsma owns eval and monitoring for it, and Brackenfen Podcast Network is one of the networks running on it."
Why this works
A "how much should eval cost" question turns into a vague percentage fast. One product turns it into a real ranking problem.
2
Say your structure out loud before naming a single number
Say it like this
"I'm going to name what the ratio is actually protecting, say which side is hardest to undo if I get it wrong, say what has to exist before what, name a cheap check I'd run before trusting either number, then state the order the budget should follow."
Why this works
Tells the interviewer you have a method for ranking a budget, not a percentage you're guessing out loud.
3
Name the outcome the ratio is actually protecting
Say it like this
"Every dollar in either line is competing to protect one thing: a show note nobody at a sponsor or a network ever has to double check. Rank the two costs without naming that first, and you're just picking a percentage that feels safe."
Why this works
Without a named outcome, "spend this much on eval" is an opinion wearing a budget line.
4
State the ratio itself, and how it should move by stage
Say it like this
"Early on, eval and monitoring should cost more than serving, for Reedline it started at eleven times. As volume scales, that ratio should fall, but it should land on a floor, not zero. Reedline's floor sits around seventeen percent of serving cost, and it stays there for as long as the product runs."
Why this works
This is the actual answer to the question. Everything else is why it's the right shape.
5
Say what's hardest to undo, and why that's what gets ranked first
Say it like this
"If I ranked by which line is bigger this month, I'd cut eval the moment serving cost caught up. But a wrong name that reaches a sponsor's compliance report can't be un-sent. Adding serving capacity back is a purchase order. That's why eval gets funded first even on the month it isn't the biggest number."
Why this works
This is the whole test of the framework. A ranking that only follows this month's dollar figure gets the order wrong.
6
Name what has to exist before what
Say it like this
"You can't safely take a network from 240 episodes to 90,000 without the golden set and the automated eval harness built first. Serving capacity is downstream of that, not the other way round."
Why this works
Naming the dependency stops a team from selling scale it hasn't actually built the safety net for yet.
7
Name the cheap evidence, then close on the stated order
Say it like this
"A small quarterly audit, twenty episodes recut by cohort instead of the blended average, is what actually caught this. So: fund the golden set and the harness first, fund targeted review for new-host episodes second, fund serving capacity third, and let the ratio fall from eleven times to about seventeen percent, never to zero."
Why this works
Ending on the stated order, defended in one line, is what makes this sound like a ranked decision instead of a percentage recited from memory.

Let's learn

Every week, a host at Brackenfen finishes recording, and by the time they've poured a coffee, Reedline has already handed back a transcript and a full draft of the show notes: guest names, timestamps, and the three quotes worth pulling for a clip.

Before Reedline, a producer wrote show notes by hand after every episode, about fifty minutes each, checking the recording back for exact quotes and spellings. That only happened for the episodes a producer had time for, so about half of Brackenfen's back catalog never got notes at all.

Reedline drafts a full set of show notes in under three minutes, for every episode, published or not. Brackenfen went from a 240-episode pilot across three shows to 14,000 episodes a month across forty networks in about nine months.

Knowledge spark: what's an eval harness? A set of automatic checks that grades the model's own output against the transcript it came from, catching a wrong name or a made-up quote before a person ever reads it.

The turn: the extra mistakes were never really about Reedline getting worse. Serving cost barely moved; the model ran the same version the whole quarter. The turn is that Briarcask's review policy said sample one in five episodes by hand, a rule written back when every host on the platform had been recording for years. It stayed at one in five as new hosts started joining faster than the model had time to get used to them, and a flat sample kept missing the exact episodes most likely to have a name wrong.

The model didn't get worse. The sample just stopped watching where the mistakes actually were.

Here's the build-up behind the floor Reedline eventually landed on. By month twenty two, four thousand five hundred and ninety dollars a month was going into eval and monitoring, and almost none of it was a person reviewing an episode nobody needed to worry about.

The build-up: Reedline's eval and monitoring budget at month 22, by part
$5k $2.5k 0 Judge $2,700 Targeted review $1,240 Drift monitors $410 Golden-set upkeep $240 Total $4,590
Automated judge, all episodesTargeted human reviewDrift and monitoringGolden-set upkeep
Almost sixty percent of the line is the automated judge, running on every episode. The one slice that's still a person is aimed at exactly the episodes most likely to need it.

By month nine, the ratio had already fallen to about one point two times serving cost, and on a spreadsheet that looked like the plan working exactly as intended. It was, mostly. The part the ratio couldn't show was which episodes that shrinking review budget was actually landing on.

Eval and monitoring cost, as a multiple of serving cost, month 1 to month 22
12x 6x 0 the quarterly audit floor, about 0.17x Mo. 1 Mo. 5 Mo. 9 Mo. 14 Mo. 22 11.0x
Eval and monitoring, as a multiple of serving cost
Eleven times serving cost at month one, down to about one point two times by the audit at month nine, settling near a floor of zero point one seven times by month twenty two. Both dollar lines kept growing the whole time. Only the ratio between them was designed to shrink.
Hand sketched comparison diagram titled which one can you turn back next quarter. Left, a gauge icon, labeled serving capacity, captioned add it, cut it, any week you like. Right, a document icon, labeled a sponsor's trust, captioned a wrong name in a compliance report does not un-send.
Serving capacity is a dial you can turn back any week. A sponsor's trust, once a wrong name reaches their compliance report, isn't. That's the whole reason eval gets ranked first even on the month it isn't the biggest number.
The choice that mattered Briarcask set human review at a flat one in five episodes back at pilot launch, when Brackenfen had three shows and every host had been recording for years. That was a fine rule then. It stopped being fine the day new hosts started joining faster than the model had episodes of their voice to learn from, and the sample kept spending its budget on hosts who didn't need it.

At its worst, a review policy nobody revisits keeps sampling the wrong slice every month, sponsors keep catching what the audit should have caught first, and a network stops trusting the auto-published notes enough to ask for full human review again, more expensive than never automating in the first place.

What I'd leave alone: established-host review at about one percent genuinely doesn't need to grow. Those episodes were running under one percent errors the whole time this was happening, and spending more there wouldn't have caught the pattern any faster.

The lesson: a ratio that's falling exactly the way you planned can still be falling for the wrong reason underneath it. Watch what the shrinking budget is actually still covering, not just how small the number got.

Now here is the same thing as a story

Read the long version below when you want to feel why a ratio falling exactly on schedule still missed something, not just be told that it did.

Esben Brandsma can listen to the first ninety seconds of a podcast episode and tell you whether Reedline is going to have an easy time with it or a hard one. He'd built eval harnesses for a speech recognition team for four years before Briarcask hired him to own Reedline's cost and quality together.

The pilot months were good, genuinely good. Three shows, 240 episodes a month, and Esben's team reviewed every single one by hand while they built the first golden set. By month five, Brackenfen wanted every show on the platform, not just the pilot three, and volume started climbing fast.

It faded in three beats, and none of them looked like a mistake at the time. Beat one, back at pilot launch, the review policy said sample one in five episodes by hand, chosen because it caught almost every error when every host had been recording for years and the model already knew their voices well. Beat two, as volume climbed past a thousand episodes a month, Esben's team built an automated judge to run on all of them, and folded the old one-in-five human sample in underneath it, unchanged, as the safety net for whatever the judge might miss. Beat three, new hosts started joining Brackenfen and the other networks faster than before, syndication deals, guest-hosted spinoffs, and each one arrived with names and terms the model had never heard, spread thin across a sample that still treated every episode the same.

It surfaced on an ordinary Tuesday, in a routine review nobody expected to find anything. Every quarter, Briarcask's compliance lead pulls twenty episodes at random and reads them cold, mostly to keep the automated judge honest. This time she also asked for the same twenty split by how long the host had been recording, since a sponsor had emailed the week before about a guest's company name being spelled wrong in an ad-verification report.

Esben pulled the numbers that afternoon. The blended error rate on the dashboard had read a steady two point two percent all quarter, healthy, nothing to flag. Split by host tenure, established-host episodes were running at zero point six percent. New-host episodes, about one in eight of everything Reedline touched, were running at thirteen point four percent, more than twenty times higher, and the blended average had been hiding it the entire time.

We weren't chasing a broken review policy. We were chasing a healthy looking number that had quietly stopped describing the episodes it was supposed to be watching.

The dollars were smaller than the pattern. Three sponsors across the quarter had already caught a wrong company name in their ad-verification reports and asked for a make-good, about four thousand two hundred dollars in credits Brackenfen had to eat before anyone traced where the errors were actually coming from.

The decision that opened the door went back to the very first review meeting, before Brackenfen had signed a single network beyond the pilot three. Someone asked whether the sample should account for how new a host was. The honest answer at the time was no, every host on the platform had been recording for years, there was no such thing as a new one yet. Nobody wrote that assumption down as an assumption. It just became the rule, and the rule outlived the reason for it.

Run that quarter again with one change: any host in their first two weeks gets full review, human and automated both, everyone else gets the same one percent spot check as before. New-host episodes take about forty minutes to clear instead of three, but there are only about seventeen hundred of them a month out of fourteen thousand. The thirteen point four percent error rate on that slice would have been caught before publish instead of after a sponsor's email, and the quarter's make-good bill drops from about four thousand two hundred dollars to close to nothing.

One design assumed the mix of hosts would stay the same as the day the policy was written. The other design checks who is actually in the sample this month, and spends the review budget there instead of wherever the rule happened to point five years ago.

What I'd tell myself, back in that first review meeting: write down the assumption behind the sample, not just the sample. A rule with no stated assumption can't tell you when it's quietly stopped being true.

ORDER, and the five calls that set the ratio

Not a story wearing a framework's clothes. This is a budget-priority problem, and ORDER is what stops "which line looks bigger this month" from quietly standing in for "which one can't be undone."

OOutcome. What is every dollar in either line actually competing to protect?
A show note nobody at a sponsor or a network ever has to double check, at a cost per episode that doesn't outgrow what a network pays Briarcask. Rank without naming that first, and the ratio is just a percentage that feels safe.
Say the outcome before naming a single dollar figure, or the ratio is opinion wearing a budget line.
RReversibility. Which decision is hardest to undo?
Cutting eval and monitoring loses this test even in the one month it costs more than serving. A wrong name reaching a sponsor's compliance report can't be un-sent, and it can cost the account. Serving capacity is a purchase order either direction.
This is the hardest step, and the one a flat percentage skips. The line that's smaller in dollars this month isn't automatically the one you can cut.
DDependency. What has to exist before what?
You can't safely carry a network from 240 episodes to 90,000 without the golden set and the automated judge built first. Serving volume is downstream of the eval harness, not the other way round.
Naming the dependency stops a sales team from promising scale the safety net was never sized for.
EEvidence. What could you learn cheaply before committing a quarter's budget?
A twenty-episode audit, recut by host tenure instead of the blended average, is what actually found the gap here, for the cost of one afternoon.
Cheap evidence beats a percentage that's never been checked against who is actually in the sample.
RRank. State the order, defend the top pick.
In order: the golden set and eval harness first, targeted review for new-host episodes second, serving capacity growth third. Eval gets funded first because it's the only line that can't be bought back once a wrong name has already shipped.
If the order would look the same with a different outcome in step one, it was ranked by gut and the outcome got written afterward.

Three things worth stating directly, since this is where the real judgment sits. The alternative Esben's team considered first, and dropped, was doubling the flat sample from one in five episodes to one in two, across every show. It lost fast: roughly eighty eight percent of that added review would have landed on established-host episodes that were already running under one percent errors, nearly doubling review cost while barely touching the slice that was actually failing. The AI-specific failure worth naming by name is cold start on an unfamiliar voice: a new host arrives with names, company terms, and a way of talking the model hasn't calibrated to yet, so hallucinated or misheard names cluster on exactly those episodes and disappear inside a blended average that mostly describes hosts the model already knows well. The guardrail is cohort-based monitoring, splitting the error rate by host tenure instead of trusting one platform-wide number, paired with an automatic elevated review tier for any host's first two weeks. That guardrail isn't free. New-host episodes now take about forty minutes to clear instead of three, a real latency cost Briarcask accepted because those episodes are a small share of volume and sponsors' turnaround windows run twenty four to forty eight hours, long enough that the extra wait never actually threatens a deadline.

And if you want to be sure it really works, try it somewhere else

Same five moves, a farming cooperative's leaf photos instead of a podcast feed, and this time the lever was a camera, not a calendar.

CropScan is Grovemont AgriTech's AI tool. A field scout photographs a crop leaf on their phone, and CropScan flags disease risk back in under ten seconds, so nobody has to wait for an agronomist to drive out and look. Rasha Odede runs cost and quality on it, and Marrowvale Growers Cooperative is one of the co-ops running it across its member farms.

The build-up: CropScan processes about 5,000 scans a month early on, at roughly five cents a scan. Eval and monitoring ran about two thousand one hundred dollars against two hundred fifty dollars of serving cost, eight point four times over, while Grovemont built its golden set of confirmed disease photos and had an agronomist review every flagged scan by hand. By month twenty, with 300,000 scans a month across the co-op's full membership, serving cost had grown to fifteen thousand dollars and eval and monitoring had settled near two thousand eight hundred fifty dollars, about zero point one nine times serving, close to the same floor Reedline found.

The decision Rasha would take back Trusting one blended accuracy number across every scout's phone, instead of checking whether a newly donated low-cost camera model was producing images CropScan had never been calibrated to read.

The fix Grovemont tried first wasn't the same one Briarcask reached for. They considered replacing every scout's phone with one standard, higher-end model, removing the image-quality variable entirely. It lost on time, not logic: fitting two full growing seasons to re-equip every scout across the co-op, while a monitoring change that flagged results by device model shipped in a week and cost almost nothing.

Same rank, different lever: for Reedline the lever was a calendar, how long a host had been recording. For CropScan it's a camera, which phone model actually took the photo. A newly donated low-cost device, sent to volunteer scouts in one district, produced images CropScan read correctly less than half the time it caught the same blight on an established phone. The blended accuracy number never dropped enough to notice, because that district was a small share of total scans, the same way new-host episodes were a small share of Reedline's.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: spend more on eval than serving while small, shrink it only by automating the harness, and never let it hit zero, because the thing you can't undo is what sets the floor.
Cost: there's no budget this quarter for both a bigger golden set and more serving capacity. Fund the golden set. A cheap eval harness on a small pile of traffic beats a fast pipeline nobody has checked against a device or a host it hasn't seen yet.
The model got better, for real: say Reedline's underlying transcription model gets meaningfully more accurate on unfamiliar voices. That's real, and it should let the review tier for a new host's first two weeks shrink over time. It doesn't mean the tier goes away, a genuinely new voice is still a genuinely new voice to whatever model comes next.

Where people run it wrong.
They treat the eval-to-serving ratio as one fixed percentage instead of something that's supposed to fall as the product matures.
They let a sampling rule that was right on day one keep running unexamined long after the mix of what it's sampling has changed.
They cut monitoring the moment its dollar figure looks big next to serving, without asking which of the two is actually the one they can't undo.

How to use it live. Say the real question out loud before naming a percentage: is the population this rule samples still the same population it was written for. That buys a beat to check instead of defending a number nobody's re-derived in months.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
ORDER: rank by what's hardest to undo. Built for prioritization questions, including how to split a budget, not a single number to estimate.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Esben Brandsma, who owns eval and monitoring for Reedline at Briarcask Audio. Spent four years building eval harnesses for a speech recognition team before this.
3 · THE OLD HABIT
What rule kept running long after the reason for it had changed?
Tap to flip
ANSWER
A flat one-in-five human review sample, set when every host had been recording for years. Nobody revisited it as new hosts started joining faster than the model could learn their voices.
4 · THE RANKING LOGIC
Why does eval and monitoring get funded first, even in the month it isn't the biggest number?
Tap to flip
ANSWER
Because a wrong name reaching a sponsor's compliance report can't be un-sent, while serving capacity can be added or cut with a purchase order either way.
5 · THE OLD DECISION
What decision would Esben take back?
Tap to flip
ANSWER
Never writing down the assumption behind the one-in-five sample, that every host had been recording for years, so nobody noticed when that assumption stopped being true.
6 · THE NUMBER
Fill in the blank: the eval-to-serving ratio started at ___ times serving cost and settled near ___ times at full scale.
Tap to flip
ANSWER
11 times, and 0.17 times. Both lines kept growing in dollars the whole time. Only the ratio between them was designed to shrink.
7 · THE REPLAY
Same quarter, new review design, what changes?
Tap to flip
ANSWER
Any host's first two weeks get full review instead of a flat sample. New-host episodes take 40 minutes instead of 3 to clear, but the quarter's sponsor make-good bill drops from about $4,200 to close to nothing.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the different lever there?
Tap to flip
ANSWER
CropScan, a crop disease scanner at Grovemont AgriTech. There the lever is which phone model took the photo, not how long a host has been recording.

Check yourself Score: 0 / 0

True or false
1. True or false: once Reedline's eval-to-serving ratio reached its floor of about 0.17 times, it would have been safe to let it keep falling as volume kept growing.
  • True
  • False
Show hint
Ask what a floor is for. Why does the priority list say the ratio should never hit zero?
Show answer
False. A wrong name reaching a sponsor's compliance report never stopped being possible. The floor exists because monitoring has to keep catching the next cold-start pattern, not just the one that already got fixed.
Multiple choice
2. Why did the blended error rate stay a healthy looking 2.2 percent while new-host episodes were already failing at 13.4 percent?
  • A. The automated judge was miscalibrated for every episode equally.
  • B. New-host episodes made up only about one in eight of total volume, so their high error rate barely moved the platform-wide blended average.
  • C. Serving cost dropped at the same time errors rose.
  • D. Established-host episodes stopped being reviewed at all.
Show hint
Work out how much a small slice of traffic can move an average, even when that slice is failing badly.
Show answer
B. About one in eight episodes at 13.4 percent, blended with the rest at 0.6 percent, lands close to 2.2 percent. The small share hid a large problem inside a healthy looking number.
Fill in the blank
3. The eval-to-serving ratio fell from ___ times serving cost at month 1 to about ___ times at month 22.
Show hint
Check the two labeled points at the start and end of the ratio line chart in Section 1.
Show answer
11 times, then 0.17 times. Both are pulled straight from the line chart, and they're the two numbers the direct answer opens and closes on.
Short answer, name the rejected alternative
4. What did Esben's team consider first to fix the review gap, and why did it lose?
Show hint
Look at the paragraph right after the five ORDER steps, where the rejected fix gets named.
Show answer
Model answer: Doubling the flat sample from one in five episodes to one in two, across every show. It lost because roughly 88 percent of that added review would have landed on established-host episodes already running under one percent errors, nearly doubling cost without fixing the actual gap.
Short answer, apply it yourself
5. Pick an AI product you use that reviews or checks its own output at some fixed rate. Name one way the population it's sampling might have changed since that rate was set, and how you'd check.
Show hint
Think of a product that expanded into a new language, a new user group, or a new kind of input after its review rate was first decided.
Show answer
Model answer: A customer support chatbot might review a fixed 10 percent of tickets, a rate set when almost every ticket came in English. If the product later expanded to more languages, errors could cluster in the newer languages while a flat 10 percent sample barely touches them. I'd recut the sampled error rate by language and check whether any one language is running far above the blended average.
Fill in the blank, work the number
6. If Reedline's new-host share of volume grew from one in eight episodes to one in four, with the same 13.4 percent new-host error rate and 0.6 percent established-host rate, the blended error rate would rise to about ___ percent.
Show hint
Blended rate equals each cohort's share times its own error rate, added together: 0.25 times 13.4, plus 0.75 times 0.6.
Show answer
About 3.8 percent. 0.25 times 13.4 is 3.35, plus 0.75 times 0.6 is 0.45, for 3.8 total. As the risky cohort grows as a share of volume, the review budget aimed at it has to grow with it, not stay fixed at whatever worked when that cohort was smaller.
Before you close the answer
Why this works
Tests whether you'll treat the eval budget as one fixed percentage or as a ratio that's supposed to move as the product matures, and whether you know which side of it you genuinely can't undo. Most candidates guess a number. They don't rank it.
Follow-up traps
"Isn't eleven times serving cost wildly wasteful for a pilot?" Response: no, the absolute dollars are small at pilot scale, $1,584 a month. Wasteful is measured in what a wrong name costs a network's sponsor relationship, not in the ratio.

"Couldn't Briarcask just review every episode by hand forever and skip the automation?" Response: that's the $27,000-a-month serving line at full scale, with a review team sized the same way. The ratio never falls, and Briarcask prices itself out of every network that isn't a pilot.
If pressed
The automated judge itself gets audited against the golden set every two weeks, not trusted as ground truth forever, because the judge model can drift the same way the underlying transcription model can, and nobody at Briarcask wants to find that out from a sponsor's inbox.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more