ConceptIntermediateQuality, Cost & Token Economics / Eval design for product teams / #3
How do you decide between automated evals and human review?
Pick automated as the default, not because human review is worse, but because the two kinds of miss cost completely different amounts to find.
The answer, first
Default to automated evals for anything with a checkable structure or a reference to score against, a matching translation, a banned phrase, a length or format rule. Save a person's time for the calls a score cannot structurally make: whether something reads wrong in a market, whether a phrase means something different under local law, or any case where the right answer is genuinely a judgment call and not a lookup. Route by category risk and by whether a real golden set exists yet, not by the score alone, because a high score only proves the check ran clean, not that it was the right check to run.
In the order that actually matters
Default to automated evals wherever there is a checkable structure, and send the rest to a person.Why: an automated only gate misses judgment calls that pass every structural check; a human only gate is too slow to run on everything, so it quietly runs on almost nothing.
Route by category risk and language pair maturity, not by the score alone.Why: a score tells you if a listing matches the reference. It never tells you if the reference covers what actually matters in that market.
Set kill criteria before an incident forces one on you.Why: no golden set yet, or a regulated category, both mean human first, no exceptions for a high score. Waiting for the first bad listing to define the rule means the rule shows up one incident too late.
Run a standing audit sample on what got auto approved, not just what got flagged.Why: the automated blind spot is narrow, but it stays invisible if nobody ever checks what quietly passed.
Leave routine, fully automated re-translations alone.Why: a template refresh with no new claim and no new category is exactly the case automated checking was built for. Adding a person there only slows a seller down for no real safety gained.
When something does slip through, show the audit that caught it.Why: that is what proves the system is watching itself, not something you got lucky on once.
How to answer this, stage by stage
Nobody is grading whether you can list reasons automation is good. They are grading whether you will commit to a pick and still tell them, unprompted, exactly when you would break it.
1
Scope it to one real product before arguing anything in the abstract
Say it like this
"Let's ground this in one product. Aravalli is a global marketplace. LinguaCheck is the tool that checks a seller's listing after it gets machine translated, before it goes live in a new country's store. Basilio Kapadia runs quality on it."
Why this works
A tradeoff question answered in the abstract turns into a debate about automation in general. One product turns it into a real decision.
2
Say your structure out loud before you start reasoning
Say it like this
"I'm going to pick a side first, not hedge. Then I'll say who gets hurt by each kind of mistake, find the real cost gap between the two, and give you the line where I would flip my own answer."
Why this works
Tells the interviewer you have a plan, and that you can commit, which is exactly what a tradeoff question is testing.
3
State the position, before any reasoning
Say it like this
"My default is automated. Score every listing against a golden set of real reference translations. If it clears the bar, it goes live. A person only sees it when the question stops being 'does this match' and becomes 'is this actually right for this market,' because a score cannot answer that one."
Why this works
A strong tradeoff answer commits in the first breath. "It depends" fails the question before the reasoning even starts.
4
Name who feels each kind of mistake
Say it like this
"If we only automate, the mistakes that get through are the ones that look clean but are wrong for the market, and the customer reading that listing feels it first, not us. If we only use people, the review queue cannot keep up with fifty thousand listings a week, so most listings quietly stop getting checked by anyone, and everyone counting on that review feels it."
Why this works
Turns an abstract tradeoff into two named groups paying two different costs, instead of one metric arguing with another metric.
5
Find the real cost asymmetry
Say it like this
"Here's the asymmetry. A gap in the automated score is narrow, and it's cheap to find, one audit sample tells you. A gap in human coverage is broad, and it hides, because a slow review queue does not shrink a little under volume, it shrinks to almost nothing, and nobody notices a queue that quietly stopped being read."
Why this works
This is the heart of the whole method. Both sides fail sometimes. Only one failure is cheap to go looking for.
6
Give the kill criteria before you're asked for them
Say it like this
"I'd flip the default to human first in two cases. One, a language pair with no golden set yet, because there's nothing to score against. Two, a listing in a regulated category, health claims, kids' products, anything with a legal claim in it, no matter how high the score comes back."
Why this works
Naming your own kill line before anyone pushes back is what separates a real decision from a stubborn one.
7
Give the cheap check that proves it, right now
Say it like this
"The cheap way to test this on any product: pull a sample of everything the automated check passed, and have a person grade it cold. If the miss rate in the risky categories is way higher than in the safe ones, you've already found your line. You just haven't drawn it yet."
Why this works
Turns a philosophical debate about automation into one audit anyone could run this week.
8
Close on the position, restated
Say it like this
"So: automated by default, because it's fast and its gaps are cheap to find. Human review only where the question stops being a match and starts being a judgment call, and routed by risk, not by convenience."
Why this works
Closing on the same line you opened with is what makes a tradeoff answer sound decided, not still being worked out live.
Let's learn
Forty linguists used to read every new listing by hand before it could go live in a new country's store. That was the job at Aravalli, a global marketplace where a seller lists a product once and it gets machine translated into twenty two languages overnight.
Aravalli's tool is called LinguaCheck. Before it scored anything, every listing went to a person, no exceptions. That was slow but it was thorough. A seller could wait eleven days for a translated listing to clear review and go live in a new market.
Then LinguaCheck started scoring listings itself. It checks a translated listing against a golden set, eight hundred real reference translations per language pair, built by linguists ahead of time. It checks the words used, banned phrases, length, and format. A listing that clears the bar goes live on its own. Most listings now clear in under two hours instead of eleven days.
Knowledge spark: what's a golden set?
A set of translations a real linguist already checked and signed off on, held back and never used to train the model. Every new listing gets compared against it. It's the ruler the score is measured with.
Cost to check 1,000 listings, automated vs. a human linguist (log scale)
Automated eval, compute cost per 1,000 listingsHuman review, linguist time per 1,000 listings
Automated cost assumes one scoring call per listing. Human cost assumes about 14 minutes of a linguist's time per listing at a fully loaded rate near 36 dollars an hour. At 50,000 listings a week, checking everything by hand is not a staffing problem you can hire your way out of, it's a cost that scales straight with volume.
Ninety eight point six percent of the listings LinguaCheck auto approved held up when a person checked them again later, on a sample pulled for audit. That sounds like the whole story. It isn't. The story is in the other one point four percent, and what kind of mistake it actually was.
One of those listings was a weight loss supplement, translated into Brazilian Portuguese. The English claim said the product helps with weight loss. The Portuguese version said almost exactly the same thing, word for word, and it matched the golden set's own wording for that phrase. LinguaCheck scored it clean. What LinguaCheck never checked, because nobody built it to, is that this exact phrase crosses a Brazilian labeling rule the golden set was never written against, since the golden set was built from listings that had only ever sold in Europe.
The listing lived for nineteen days. Nobody caught it from a customer complaint. A quarterly legal sweep across health category listings found it, the same kind of check that should have run on day one, not day nineteen.
LinguaCheck's own score never moved that quarter. What moved was how many days a wrong health claim got to stay live before anyone read it again.
The choice that mattered
When LinguaCheck launched, the team set one rule: any listing that clears the score, in any category, goes live on its own, the moment its language pair has a golden set at all. That made sense the year Aravalli only sold clothes, electronics, and home goods, categories where a translation that matches is basically always a translation that's fine. It stopped making sense the year Aravalli added health products, children's items, and financial add ons at checkout, categories where a translation can match perfectly and still be wrong for the market it landed in.
What I would leave alone: when a seller updates a size chart or restocks the same shirt in a new color, using a template that already passed review once, that stays fully automated. No new claim, no new category, nothing for a person to actually judge. Adding a review step there would only slow a seller's restock down for no real safety gained.
The lesson: a passing score tells you the check you built was satisfied. It never tells you if you built the right check for that listing in the first place.
Now here is the same thing as a story
Read the long version below when you want to feel why nineteen days matters, not just be told that it does.
Basilio Kapadia has run quality on translated listings at Aravalli for four years. Before LinguaCheck, he could spot a bad machine translation by the time he finished its first sentence. Wrong words in the wrong order, a size chart that skipped a size, a measurement in the wrong unit. He caught things.
LinguaCheck came out of his own team's backlog problem. Eleven days was too long. He helped design the scoring rule himself: build a golden set for a language pair, score every listing against it, auto approve above the bar. For the first year, it worked exactly the way it was supposed to. Sellers stopped waiting. His team stopped drowning. He watched the backlog fall from eleven days to under two hours and felt, for the first time in years, like the job was sized right.
Aravalli kept growing. Clothes, electronics, and home goods, the categories LinguaCheck was built and tested on, stayed exactly as safe as they'd always been. Then the catalog team added supplements. Then children's toys. Then a small line of financial add ons at checkout, extended warranties, buy now pay later. Each addition went through the same launch checklist as everything else: build the golden set, hit the bar, ship it live on its own.
Nobody sat down and decided a weight loss claim deserved the same scoring rule as a t shirt color. It just already applied, because the rule had never been written to ask the question.
The audit found it on a Tuesday in March. Legal ran its usual quarterly sweep across health category listings, the kind of check that happens four times a year, and pulled the Brazilian Portuguese listing that had been live, and selling, for nineteen days.
Basilio didn't ask his team to retrain the score that afternoon. He asked for the last ninety days of every health category listing that had auto approved, all four hundred and twelve of them, not just the one the audit had happened to catch.
That afternoon didn't find a broken model. It found a rule that had never once been asked to know the difference between a shirt and a supplement.
The decision that opened the door went back to the meeting where LinguaCheck first launched. Someone on his team had asked whether category should change how a listing gets routed, or whether the score alone should decide. The answer, at the time, was the score alone, because category didn't change anything back then. Nobody came back to that answer as the catalog grew. Only the volume did.
Run the same Tuesday again with one change: any listing in a regulated category goes to a person, no matter what the score says, the moment it's created, not the moment a quarterly sweep gets around to it. The Brazilian listing still gets machine translated the exact same way. It still scores clean against the golden set. But it never goes live on its own, because health is on the list. A linguist reads it in eleven minutes on day one, flags the claim, and it never reaches a single customer.
One design trusted one number to answer every category's question. The other asks the risky categories a harder question than the safe ones, on purpose.
What I'd tell myself, back in that first launch meeting: the moment a rule gets applied to a case nobody actually tested it against, ask whether the rule still fits, or whether it only ever fit because that case hadn't shown up yet. Nobody asked. That's on the room, not on the score.
Four letters, since one score cannot do two jobs
This isn't a coin flip between automating everything and reviewing everything by hand. It's PICK, run until the asymmetry actually shows itself.
PPosition. Say the pick before the reasoning.
Automated by default. A person only sees a listing when the question stops being "does this match" and starts being "is this actually right here."
Say this first, or the rest sounds like you're still deciding live, in the interview.
IImpact. Who feels each kind of mistake.
Automated only: the customer reading a wrong claim, weeks before anyone checks it again. Human only: everyone counting on a review that a shrinking queue quietly stopped running.
Name both people by name, or the tradeoff stays a metric fighting another metric instead of a real decision.
CCost asymmetry. Which failure is cheap to find.
A gap in the automated score is narrow, one audit sample finds it. A gap in human coverage is broad, and it hides, because a slow queue doesn't shrink a little under enough volume, it shrinks to almost nothing, quietly, with nobody noticing.
This is the hard letter. Both sides fail sometimes. Only one failure announces itself.
Both options miss sometimes. Only one of those misses is cheap to go looking for.
KKill criteria. What flips the pick.
No golden set yet for that language pair, human first, there's nothing to score against. A regulated category, health claims, kids' products, legal language, human first, no matter the score.
Say your own kill line before anyone asks for it. That's what turns a preference into a decision.
Real miss rate on auto approved listings, by category risk score
Real miss rate found on later audit, by category risk score
Apparel and electronics sit under 0.5 percent. Health claims and children's products sit above 3 percent, six times higher, on listings that all cleared the exact same automated bar. The kill line at risk score 6 is where the team now routes to a person first, regardless of score.
Three things worth saying directly, since this is where the real judgment sits. The alternative Basilio's team rejected right after the audit was reviewing every health category listing by hand going forward, no scoring at all. It lost, because health is a sliver of Aravalli's fifty thousand weekly listings, and sending every clothing and electronics listing through the same slow queue just to catch one category's risk would have brought back the eleven day wait for sellers who never had a compliance problem to begin with. The AI specific failure worth naming is a score that measures whether a translation matches a reference, quietly standing in for whether the claim itself is allowed, two different questions that happen to look identical on a passing score. The fix that catches it is the category risk routing itself, plus a standing audit sample pulled from what LinguaCheck approved, not just what it flagged, since a blind spot you never sample for stays invisible by definition. That fix isn't free. Routing every regulated category listing to a person costs real linguist time, about four hundred listings a week, at eleven minutes each, time that used to go toward the general review queue. It's spent on purpose, only where a miss costs more than a person's morning. And the bar was never zero misses, a marketplace running fifty thousand translated listings a week on a probabilistic model cannot promise that. It's an audited miss rate, checked monthly against a live sample, not a number assumed safe because the launch review once said so.
And if you want to be sure it really works, try it somewhere else
Same four letters, an appliance factory floor instead of a translation queue, nothing about e commerce anywhere in sight.
SentryLine is a camera system at Ostrand that checks welded parts on a dishwasher assembly line for defects. Draven Vasari is the quality engineer who decides how flagged parts get handled.
The case for automating the whole check: with six inspectors covering three shifts, only about one in twelve control panels ever got a full human look. A wiring defect on a gas line bracket went unseen for weeks and reached customers before a recall caught it.
The case against automating the whole check: an earlier build of SentryLine auto rejected or auto passed every part purely on its own confidence score, safety rated parts included. It auto passed a bad weld on that same kind of gas line bracket, because the defect was a pattern the model had never seen before, and a score with nothing to compare it against just guessed confident and wrong.
The decision Draven would take back
Routing every part by confidence score alone, instead of pulling every safety rated part class onto a human first list from the start, no matter how sure the model sounded.
The score alone cannot tell a routine bracket from a gas line bracket. The part list has to do that job instead.
Same rank as before: automated by default, because most of what the camera checks is a real match to a known good pattern. Human first only where a wrong call is not something the factory can afford to learn about later, which is exactly the safety rated list, decided ahead of time, not discovered after a recall.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the pick, automated by default, human first only on the kill list, say the kill list out loud and stop.
Cost: there's only budget this quarter for one project, a faster model or a bigger review team. The kill list wins every time, a faster wrong answer on a regulated listing isn't worth having.
The model got better, for real: say the overall pass rate climbed this quarter. That isn't proof the regulated categories climbed with it. A model can look better on average while the categories that cost the most stay exactly as blind as before.
Where people run it wrong.
They read one overall pass rate and assume it covers every category equally.
They treat "we don't have the headcount for full human review" as a reason to skip routing by risk, instead of the reason routing by risk matters more.
They wait for the first bad case to define the kill criteria, instead of naming the criteria before anything has gone wrong yet.
How to use it live. Say the real question out loud before answering it: "is this asking me to pick automated or human, or asking whether I can tell the difference between a check a score can make and one only a person can." That buys you a breath, and it tells the interviewer you already see past the trap.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
PICK: commit to a position first, then show the real cost gap between the two ways a decision can fail. Built for tradeoff questions, "A or B."
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Basilio Kapadia, quality lead for translated listings at Aravalli, a global marketplace. Helped design the scoring tool, LinguaCheck, himself.
3 · THE POSITION
What's the position, stated first, before any reasoning?
Tap to flip
ANSWER
Automated evals by default. A person only reviews a listing when the question is a genuine judgment call a score cannot make.
4 · THE COST ASYMMETRY
What's the real cost gap between the two kinds of miss?
Tap to flip
ANSWER
Automated blind spots are narrow and cheap to find with one audit sample. Human coverage gaps are broad and stay invisible, because a slow queue shrinks to almost nothing under volume, quietly.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Routing every listing by score alone, in every category, once a language pair had a golden set at all, instead of routing by category risk from the start.
6 · THE NUMBER
Fill in the blank: 98.6 percent of auto approved listings held up on audit. The Brazilian listing that didn't lived for ___ days before a quarterly sweep caught it.
Tap to flip
ANSWER
Nineteen days. The blended score barely moved, but that's exactly the number that shows why a blended score hides the real risk.
7 · THE REPLAY
Same bad case, new design, what changes?
Tap to flip
ANSWER
Category risk routing sends the health listing to a person on day one, not a quarterly sweep. A linguist reads it in eleven minutes and it never reaches a customer.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the same tension?
Tap to flip
ANSWER
SentryLine, a defect camera at appliance maker Ostrand. Same tension: routing purely by confidence score misses the same kind of case, a safety rated part the model has never seen before.
Check yourself Score: 0 / 0
Fill in the blank
1. Before LinguaCheck, the review backlog ran about ___ days behind. After, most listings clear in under ___.
Show hint
It's the two numbers named right at the start of the "Let's learn" section.
Show answer
Eleven days, then under two hours. That's the gap a fully human process couldn't close at fifty thousand listings a week.
Multiple choice
2. Why can't Aravalli just fix this by running "a bit more" human review across the board, instead of splitting listings by category risk?
A. Because linguists refuse to review more than a set number of listings a day.
B. Because spreading more review evenly across every category still doesn't tell you which listings actually needed a person in the first place.
C. Because the model would need to be retrained before any more review could happen.
D. Because customers never read the listings closely enough for review to matter.
Show hint
Think about what the category risk chart shows: the miss rate isn't spread evenly across categories.
Show answer
B. "A bit more review, everywhere" is a dial, not a decision. The real fix targets the categories where a miss actually costs something, not every category equally.
True or false
3. True or false: because LinguaCheck's overall auto approve rate held up at 98.6 percent, Aravalli could safely spend the same review budget on every category.
True
False
Show hint
Check the category risk chart. The blended number and the by-category numbers tell two different stories.
Show answer
False. The 98.6 percent was an average across categories with very different real risk. Health and children's categories ran a miss rate six times higher than apparel and electronics, hidden completely by the blended number.
Short answer, name the rejected alternative
4. What alternative did Basilio's team consider right after the audit, and why did it lose?
Show hint
Look at what "half the room" would have wanted after finding a real miss, and what the framework recap paragraph says lost instead.
Show answer
Model answer: Reviewing every health category listing by hand going forward, with no automated scoring at all. It lost because health is a small share of the fifty thousand weekly listings, and sending every clothing and electronics listing through the same slow queue to catch one category's risk would have brought back the eleven day wait for sellers who never had a compliance problem.
Short answer, apply it yourself
5. Pick an AI product you use yourself. Name one place its automated check is structurally unable to catch a real problem, and how you'd check for it.
Show hint
Think of a product that scores or ranks something against a fixed reference, where the reference itself might not cover every real case.
Show answer
Model answer: A resume screening tool that matches keywords to a job description can score two candidates as an equal match, while one genuinely has the seniority the role needs and the other only has the words. I'd check by pulling a sample of borderline-score resumes each month and having a recruiter grade them by hand, instead of trusting the match score alone.
Multiple choice
6. A brand new language pair launches with zero golden set translations built yet. What do the kill criteria say to do?
A. Auto approve using the closest related language's golden set until a real one is built.
B. Route every listing in that language pair to a person, regardless of any score, until a real golden set exists.
C. Wait for enough volume to build up before deciding how to review it.
D. Skip quality checks for that language pair until the golden set is ready.
Show hint
Check the K step: it names "no golden set yet" as one of exactly two kill conditions.
Show answer
B. A score with nothing real to compare against isn't a low confidence score, it's not a real score at all. Human review is the only option that actually exists until a golden set is built.
Before you close the answer
Why this works
Tests whether you can commit to a default and still hold a real line for when you'd break it, instead of answering "it depends" or defending one option as always right.
Follow-up traps
"Doesn't a 98.6 percent pass rate already prove automated review is good enough everywhere?" Response: that number is a blend across categories that don't carry the same risk. The miss rate in regulated categories alone ran about six times higher, and that's the number that actually decides the routing.
"If human review is too slow to run on everything, why not just hire more linguists?" Response: the problem isn't headcount, it's that any fixed review rate becomes a smaller and smaller share of a growing volume. Doubling the team shrinks the backlog for a while. It doesn't remove the deeper problem.
If pressed
The category risk score itself isn't a one time judgment call. It's recomputed monthly from the count of real compliance flags and legal takedowns filed against that category in the last 90 days, so a category can move in or out of the human first tier as real incidents build up, not by someone's opinion at launch.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.