ConceptAdvancedEval-Driven Specification / Writing an eval spec / #24

Explain how an eval spec should handle cost and latency alongside quality.

The direct answer
Write the spec so a call only passes when three limits hold at once: a quality floor, a cost ceiling, and a latency ceiling, each tied to how the score actually gets used. Set every number against something real, not a guess: human-rater agreement for quality, per-call revenue for cost, how fast a person would notice on their own for latency. Then say plainly that cost is the one allowed to bend first, and name exactly what it can never take down with it, like the step that keeps the two speakers on a call apart.
Do this, in order
  1. Set three limits that all have to pass together: a quality floor, a cost ceiling, a latency ceiling, not one number standing in for the whole spec.Why: a call center trusts the score for coaching and for an agent's file, so a cheap, fast, wrong answer is worse than no answer.
  2. Break the cost side into its real parts before picking a ceiling: transcription, the model reading the call, the model writing the score back out.Why: each part moves differently as the call gets longer or the mode changes, and a target with no build-up under it is a guess wearing a number's clothes.
  3. Give every limit a low and a high tied to how the score gets used, not to the call's length.Why: a live coaching alert and an overnight QA record are graded on completely different clocks.
  4. Check each number against something outside the spec: human agreement for quality, revenue for cost, a person's own reaction time for latency.Why: a target with no outside anchor is a guess with a decimal point.
  5. Name cost as the limit most likely to bend first, and say what it can never cut to get there.Why: cost is the only one of the three that can fail quietly, so the spec has to draw its own line, not leave it to whoever's under budget pressure that month.
  6. Re-check the whole budget the moment call volume or average call length shifts.Why: skip this and the numbers stay accurate for a call center that no longer exists.

How to answer this, stage by stage

Nobody is grading whether you land on exactly nine cents. They are grading whether all three limits are real, checked against something outside the spec, with an honest answer for which one gives first. Seven moves get you there.

1
Pin the call down as the unit, and name the three limits before any numbers
Say it like this
"Before I put a number on anything: one call here is one finished support call, scored either after it ends or live while it's still running, those are two different jobs. And this isn't one target. It's three limits that all have to hold on the same call: how good the score has to be, what it's allowed to cost, and how fast it has to land."
Why this works
Frames the spec as three limits that hold together, not one blended metric, before the interviewer can assume it's a single number.
2
Say the pass or fail rule out loud
Say it like this
"Here's the rule. A call passes the spec only if the score clears the quality floor, the cost lands under the ceiling, and the answer comes back before the latency ceiling. Miss any one of the three and the call fails, even if the other two look great."
Why this works
States the method before a single number lands, so what follows reads as arithmetic against a stated rule, not numbers pulled from nowhere.
3
Own the numbers for one real call
Say it like this
"Take an average seven-minute support call, scored once, overnight. Transcription runs about seven tenths of a cent a minute, so that's 4.9 cents. The model reading the transcript plus the scoring instructions is about fifteen hundred tokens, under half a cent. Writing the score and the two quoted reasons back out is under a fifth of a cent. Add it up: about five and a half cents to score one call, once."
Why this works
Turns "a cost ceiling" into an actual number a reviewer can check line by line.
4
Split the range by how the score gets used, not by how long the call runs
Say it like this
"That five and a half cents is the overnight number, one pass, ready for tomorrow's coaching huddle within six hours. Live coaching is a different job. The model has to re-check the call about once a minute while it's happening, so a seven-minute call gets scored roughly seven times, more like nine cents. The quality bar moves too: the overnight score goes in someone's file, so I'd hold that to ninety percent agreement with a senior reviewer. The live alert is just a nudge with a supervisor still listening, so eighty-two percent is enough, and it has to land within twenty seconds, not six hours."
Why this works
Shows all three limits move together depending on how the number gets used, not just cost against a single axis like call length.
5
Check every number against something outside the spec
Say it like this
"Two senior QA reviewers scoring the same calls only agree with each other about ninety percent of the time. So a ninety percent floor isn't strict, it's about where human agreement itself tops out. On cost, the call center pays about thirty-five cents a call under the contract, so five or nine cents leaves real room. On latency, a supervisor listening live usually notices trouble on their own in about forty-five seconds. Beat that by more than half and the tool is actually adding something, not just repeating what a person would have caught anyway."
Why this works
This is the step most estimates skip, the one that turns a plausible-sounding number into one that survives a follow-up question.
6
Name which limit bends first, and why nobody notices right away
Say it like this
"If I had to guess which one gets quietly traded, it's cost. Quality failing gets noticed the day someone disputes a score. Latency failing gets noticed the moment an alert shows up after the call already ended. Cost failing doesn't trip anything on its own, someone swaps a cheaper vendor or trims a step, the bill goes down, and nothing on the screen changes that day. So the spec has to name what cost is never allowed to cut, not just what it's allowed to spend."
Why this works
Answers the hardest part of the question directly, and shows judgment about a real production risk instead of just arithmetic.
7
Close on the rule, in one breath
Say it like this
"So: three limits, each checked against something real outside the spec, not against each other. Cost is the one I'd expect to bend first, which is exactly why the spec has to name the one piece of it, keeping the two speakers on a call apart, that can never be the thing that gets cut to hit the number."
Why this works
Restates the whole decision in one breath, the way a strong close actually sounds.
If you remember one thing A ceiling met on paper can still hide a broken call. Cost can go down without a single number in the spec moving, if what got cut inside it was never named.

Let's learn

Callridge is a tool call centers plug into their phone system. It listens to every support call and scores how the customer felt, angry, neutral, happy, so a supervisor doesn't have to sit through the whole call to find out.

Knowledge spark: what does an eval spec actually gate? It's the written rule that says whether a version of the AI is allowed to ship. Not "does it seem good," a real bar: this score, this cost, this speed, all three, checked before anyone flips it on for real customers.

Before Callridge, a supervisor at a forty-agent floor could only sit and listen to a handful of calls by hand. About two per agent a month, roughly eighty calls out of nine thousand that came in, fifteen minutes each to do it properly.

Now every call gets scored the moment it ends. A live add-on goes further: while an agent is still on the phone, if the sentiment turns negative, a supervisor's screen lights up inside twenty seconds, while the call is still running.

The extra coverage isn't what makes this risky. What makes it risky is a spec that asks for three things at once and never says which one is allowed to give first when a team can't hit all three.

At its worst, someone hits the cost ceiling by cutting the one part of the pipeline that tells the model which voice is the agent and which is the customer. The dollar number stays fine. The score is now scoring the wrong person, and nothing in the spec catches it, because nothing in the spec ever named that part as protected.

The decision that mattered Name which limit is allowed to bend, and name what inside it can never move. A ceiling with no protected parts invites exactly the cut that breaks the product while staying under budget.

The choice I would take back. We wrote the three limits with no order and no protected parts. I'd name cost as the one that bends first, and name the one thing inside it that never gets cut: keeping the two speakers apart. A cheaper transcription vendor is fine. A transcript that can't tell who's talking is not.

What I would leave alone. The little coaching-tip line the tool suggests to a supervisor, something like "try acknowledging the wait time first," doesn't need the same floor. It's a suggestion, not a record. If it's ordinary two days out of ten, nobody's file is touched by it.

The lesson. Three numbers with no order are really three teams each assuming someone else is holding the line. Someone has to write down which one bends, and what it can't take down with it, before the quarter where compute costs money nobody budgeted for.

Now here is the same thing as a story

Skip this part if you already believe a cost ceiling can be met on paper while the product quietly breaks underneath it. Read on if you don't.

Samira Haddad wrote Callridge's first eval spec two years ago, back when the company was four people pitching one regional call center. In those early months she sat in on the customer calls herself and listened, by ear, for the moments the model's score and her own read of a call didn't match.

A one-page eval spec, before and after. Before, three numbers sit with no order and nothing protected. After, the cost line is tagged bends first, with a box around keep speakers separated, never cut.
The same spec page, before and after Samira wrote down what cost couldn't touch

After the funding round, Callridge grew to sixty call-center customers in a year. For the first stretch, every model update still went through Samira's full review, including a manual listen-through of ten calls where she checked that agent and customer voices stayed cleanly split apart.

Then the review quietly thinned. An engineer took over the compute budget, and ten calls became five. Five became "spot check if something looks off." By the quarter live coaching's per-call cost ran thirty percent over its own ceiling, the listen-through hadn't happened in months, because cost fixes had stopped being something her team reviewed at all. They'd become finance's problem to solve, quietly, on their own.

The fix that quarter was a cheaper transcription vendor. It didn't include the add-on that separates the two voices on a call. Nobody flagged it, because the dollar number came in fine, and dollar numbers were the only thing anyone was still watching.

The near miss came on an ordinary Tuesday. A shift supervisor pulled up a live-coached call to write up an agent for sounding curt with a customer. The agent asked to hear the recording first. The "angry" voice tagged to them was the customer, raised near the end of a long hold. The agent hadn't said a harsh word the whole call.

We didn't lose the number. We lost which voice was which, and the number never knew.

Samira pulled the pipeline logs that afternoon. The diarization step had been off, quietly, on every live-coached call for three months. The cost line had never once crossed its ceiling. The transcript had crossed a line the spec never drew.

She remembered the meeting where the original spec got written, four people around a laptop, three numbers on a slide: quality, cost, latency. Someone said, only half joking, "if we ever have to choose, obviously it's cost." Nobody wrote that down. It felt too obvious to need a sentence.

She rewrote the spec that week. Cost stays the limit that bends first, in writing this time. But one line inside it is boxed off: whatever changes to hit the cost ceiling, the two speakers on a call stay separated, always. A monthly check now listens to ten calls by hand and compares the pipeline's own speaker labels against what a human hears. Twenty minutes, once a month.

In the next two quarters, that check caught one more vendor swap before it shipped, the same thirty percent in savings, this time with the speaker split still intact.

The thing I'd go back and tell myself, in that four-person meeting: writing three numbers down isn't the same as writing down what happens the day we can only hit two of them. I left that page blank and called it obvious.

B O U N D, mapped onto three ceilings at once

This is a sizing question with a priority problem sitting inside it, so BOUND fits, and FLIPS doesn't. Nobody's trust is flipping here. It's a budget with three lines, and an honest answer for which one moves first.

B, break it down. A call only passes the spec if it clears all three limits at once: the score at or above the quality floor, the cost at or under the ceiling, the answer back before the latency ceiling.
O, own the numbers. Quality: 90% agreement overnight, 82% live. Cost: about 5.5 cents overnight, about 9 cents live. Latency: six hours overnight, twenty seconds live. Six real numbers, not one.
U, use a range. The range comes from how the score gets used, not from how long the call ran. An overnight record and a live nudge get genuinely different numbers on all three limits, because different people are relying on them in different ways.
N, nail the sanity check. Ninety percent checked against ninety percent human-to-human agreement. Cost checked against the thirty-five cents the call center actually pays. Latency checked against the forty-five seconds a supervisor notices trouble on their own.
D, direction. Cost bends first, because it's the only one of the three whose failure doesn't trip anything on its own. That's the whole reversal: name what cost can never cut on its way down.

The build-up: what one overnight call costs to score
Transcription (7 min × $0.007)
$0.049
$0.049 of $0.0547
Model input tokens (~1,550 tok)
 
$0.0039 of $0.0547
Model output tokens (~180 tok)
 
$0.0018 of $0.0547
Total, this call
$0.0547
rounds to $0.055
Transcription Model input Model output
Transcription alone is about 90 cents of every dollar this call costs to score. That's why the transcript, and what's cut inside it to save money, matters more than which model reads it.
A number line marking cost per call scored, low to high: overnight actual 5.5 cents, overnight ceiling 8 cents, live actual 8.9 cents, live ceiling 13 cents, and what the call center pays, 35 cents, marked separately for scale.
The range, with what the call center pays marked for scale, not as one of the bounds
What moves the cost estimate most
Mode switches from overnight (1 pass) to live (~7 passes)+$0.034
Average call length grows from 7 to 11 minutes+$0.030
Diarization turned off to cut transcription spend−$0.012
Coaching-tip text roughly doubles the model's output+$0.002
Mode and call length swing the number more than diarization would save. It's on this chart anyway, because it's the one line here that can never move, no matter how tempting the savings look, since the whole score depends on it.

And if you want to be sure it really works, try it somewhere else

A wholesale produce packer runs an AI camera over the belt that sorts avocados, grading ripeness and bruising as each one passes, before a mechanical arm diverts it to export-grade or local-grade.

B, break it down. An item passes the spec only if the grade clears the accuracy floor, the compute cost per item is under the ceiling, and the sort decision comes back before the diverter physically needs it.
O, own the numbers. Accuracy floor: 92% match to a certified human grader. Cost: about $0.004 per item on the full-speed line, ceiling $0.006. Latency: 400 milliseconds, a hard limit set by how fast the belt moves the fruit to the diverter.
U, use a range. The full-speed export line grades every item at that 400-millisecond ceiling. A slower compliance-audit line re-checks a sampled fraction with a bigger model, allowed 2 seconds and a $0.02 ceiling, since it only runs on a fraction of the fruit.
N, nail the sanity check. 92% checked against 94% agreement between two certified human graders, the same shape of check as Callridge, new numbers. Cost checked against about $1.80 of margin per case, both ceilings are close to free against that. Latency checked against the belt itself, not a business tolerance, a mechanical deadline that doesn't move.
D, direction. Quality bends first here, not cost. Latency can't move, it's set by the belt's physics. Cost is close to free per item. That leaves quality as the only lever standing, and it's the one that quietly slips when a model gets tuned to always beat the 400-millisecond deadline.

Which limit bends depends on what's actually fixed At Callridge, cost bends first, because a quiet cut never trips anything on its own. On the packing line, latency can't bend at all, it's physics, and cost is nearly free, so quality is the only lever left. Look at what's actually fixed before guessing which number gives first.

Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the rule and the six numbers, five and a half cents overnight, nine cents live, checked against ninety percent human agreement and thirty-five cents of revenue. The build-up backs it up if they ask.
Cost: Callridge's finance team caps the whole live-coaching compute budget at a fixed monthly figure instead of asking what one call should cost. Same rule, solved backward: divide the budget by expected call volume and the live-versus-overnight mix to find what ceiling the numbers can actually afford.
The model got better: a new scoring model ships that's both cheaper and faster. The spec doesn't need a rewrite, just the same three checks re-run with the new price and the new speed in place of the old ones.

Where people run it wrong.
They write one flat number that quietly averages two very different jobs, live and overnight, instead of splitting the range by how the score is used.
They set a cost ceiling with no note on what's protected inside it, so any cut that keeps the dollar figure under the line looks fine on paper.
They pick a sanity check that isn't actually external, comparing the quality floor to last quarter's own score instead of to something outside the system, like how often two human reviewers agree with each other.

How to use it live. Say the pass or fail rule before any number: "this only counts as passing if quality, cost, and latency all clear their bar on the same call, not on average." That sentence buys the time to work out where the three numbers should land instead of guessing one that sounds safe.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What framework fits a question about handling cost and latency alongside quality, and why not FLIPS?
Tap to flip
ANSWER
BOUND. This is a sizing question about three numbers that all have to hold at once, not a person's trust flipping in two settings.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Samira Haddad, eval lead at Callridge, a company that sells call-scoring AI to call centers. She wrote the original eval spec two years ago.
3 · THE CHECK THAT THINNED
What check used to run on every model update, and how did it quietly shrink?
Tap to flip
ANSWER
A manual listen-through of ten calls checking that agent and customer voices stayed separated. It shrank to five, then to "only if something looks off," then stopped once cost fixes became finance's problem alone.
4 · WHAT PASSED WHILE BROKEN
What cleared the cost ceiling and looked fine on the quality score, while actually broken?
Tap to flip
ANSWER
A live-coached call with no speaker separation. The dollar cost stayed under the ceiling. The score kept looking normal, because nothing in the spec checked which voice it was actually attached to.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at first?
Tap to flip
ANSWER
Writing three limits with no order and no protected parts. It made sense when Callridge was four people and everyone assumed cost was the one you'd sacrifice, so nobody wrote it down.
6 · THE NUMBER
Fill in the blank: two senior QA reviewers scoring the same calls only agree with each other about ___ percent of the time, which is why the overnight quality floor was set at 90, not higher.
Tap to flip
ANSWER
About 90 percent. Setting the floor above that would demand more consistency from the model than two trained people manage with each other.
7 · THE REPLAY
Same near miss, new spec. What changes?
Tap to flip
ANSWER
A monthly check listens to ten calls and compares the pipeline's speaker labels to what a human hears, twenty minutes a month. In the next two quarters it caught one more vendor swap before it shipped, with the speaker split still intact.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and which limit bends first there?
Tap to flip
ANSWER
A produce-grading camera on a packing line. Quality bends first there, because latency is fixed by the belt's physics and cost is nearly free per item, unlike Callridge, where cost is the one that bends.

Check yourself Score: 0 / 0

Short answer, the number question
1. If Callridge's average call length grows from 7 minutes to 11 minutes, and every price stays the same, what's the new overnight cost per call, and does the 8-cent ceiling survive it?
Show hint
Recompute all three terms at 11 minutes: transcription at $0.007/minute, transcript tokens at about 150 per minute plus the 500-token instructions, and the score output.
Show answer
About 8.4 cents, and no, the 8-cent ceiling doesn't survive it. Transcription: 11 × $0.007 = $0.077. Input tokens: (11 × 150) + 500 = 2,150, × $2.50/million = $0.0054. Output: about $0.0018. Total is about $0.084, a bit over the 8-cent ceiling on a call that's still a normal, longer support call, not an outlier.
Multiple choice
2. Why does BOUND fit a question about handling cost and latency alongside quality, rather than FLIPS?
  • A. Because cost questions are always technical questions.
  • B. Because this is a sizing question, three numbers that all have to hold at once, with no person's trust flipping in two settings.
  • C. Because BOUND and FLIPS always get used together on eval-spec questions.
  • D. Because FLIPS only works for questions about a single call center.
Show hint
Ask what FLIPS actually needs to work: someone's habit fading, then a switch that snaps between two settings.
Show answer
B. Samira's story is about a budget built wrong, not a habit that quietly stopped working. That's exactly the shape BOUND was built for.
True or false
3. True or false: in the Callridge story, quality is the limit most likely to bend first when the three can't all be hit.
  • True
  • False
Show hint
Think about which failure gets noticed the same day, and which one can hide for months.
Show answer
False. Cost is the one that bends first. A quality miss gets disputed and a late latency alert is obvious immediately, but a cost cut can hide inside a vendor swap and never trip anything on its own.
Fill in the blank
4. The overnight build-up for a 7-minute call is about ___ cents of transcription, plus under half a cent of model input, plus under a fifth of a cent of model output, for a total of about ___ cents.
Show hint
Check the build-up chart's rows and its total row.
Show answer
4.9 cents transcription, about 5.5 cents total. $0.049 + $0.0039 + $0.0018 = $0.0547, which rounds to about 5.5 cents.
Short answer, apply it yourself
5. Pick an AI feature you use that has to be both fast and cheap to be useful, a maps app rerouting you, a phone keyboard's autocomplete. Name one limit among quality, cost, and latency that you think would bend first if the team building it had to cut something, and why.
Show hint
Think about which failure you'd notice the moment it happened, versus which one could slip for weeks without you ever knowing.
Show answer
Model answer: "A maps app rerouting me: a slow or clearly wrong route gets noticed the moment it happens, mid-drive. Cost is the quiet one. The app could quietly switch to a cheaper, slightly staler traffic feed to save money, and most drivers would never notice for weeks, the same shape as Callridge's diarization cut." Any answer works if it names a real usage pattern and explains which failure stays hidden longest.
Multiple choice
6. Why was the overnight quality floor set at 90 percent, and not 99 or 100?
  • A. Because 90 percent is a round number that looked clean in the spec.
  • B. Because two senior human reviewers only agree with each other about 90 percent of the time, so a higher floor would demand more consistency than people manage on their own.
  • C. Because the model's own confidence score reported 90 percent.
  • D. Because the call center asked for exactly 90 percent in their contract.
Show hint
Look back at stage 5's sanity check. The number wasn't picked, it was matched to something.
Show answer
B. The floor was set against the ceiling of what human agreement itself looks like, not against a feeling of what sounded strict enough.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more