InterviewIntermediateQuality, Cost & Token Economics / Eval design for product teams / #22
What eval would you build first with only one week?
With five real days, build the golden set that catches a wrongly assigned action item, not the eval that polishes a sentence someone was going to reread anyway.
The direct answer
With only a week, I would build one thing: a golden set of thirty to forty real call transcripts, graded by someone who was actually on the call, scoring whether each action item has the right task, the right owner, and the right deadline. Not a broad eval across everything the tool writes, just that one output, because it is the one nobody rereads before acting on it, and a wrong one goes unnoticed until the missed task surfaces weeks later.
Do this, in order
Build a thirty to forty transcript golden set scoring action item correctness, task, owner, and deadline, first.Why: it is the output people act on without a second look, so a wrong one costs a missed commitment nobody notices for weeks.
Use real, recent transcripts, not clean written up ones.Why: real calls carry the overlapping speech and vague phrasing that actually cause a task to land on the wrong person. A tidy fake transcript will not reproduce that.
Score invented tasks as their own category, separate from wrong owner.Why: a task nobody agreed to and a task sent to the wrong person are different failures with different fixes. Blending them into one score hides which one to fix first.
Get one real number by day three, even a rough one.Why: a week is too short to wait for a clean eval. An early number tells you by midweek whether you are anywhere near safe enough to remove the human read.
Leave the summary's wording and tone for week two.Why: a clunky sentence gets noticed and mentally fixed the moment someone reads it. It does not need the same rigor in week one.
Write down the score you need to see before you look at the number, not after.Why: setting the bar after seeing the score means the number always looks good enough.
How to answer this, stage by stage
Nobody is grading whether you know what a golden set is. They are grading whether you can pick the one mistake worth a scarce week, and defend leaving the rest for later.
1
Scope it to one real product before answering in the abstract
Say it like this
"Let's ground this in one product. Notchbrief listens in on client calls at Fenshaw and Corbeau, a small consulting firm, and about ninety seconds after the call ends it sends out a summary and a list of who owes what by when. Beckett Prentice runs eval on it, for the first time, with one week on the clock."
Why this works
An abstract "what eval would you build" answer turns into a shopping list fast. One product forces a real ranking.
2
Say your ranking rule out loud before diving in
Say it like this
"I'll answer this with one rule: rank by what's hardest to undo. So first I'll say what everything is competing to protect, then which mistake is the expensive one, then the single eval I'd spend the week on."
Why this works
Tells the interviewer you have a method, not just a list of eval ideas you happened to think of.
3
Reframe the question before answering it
Say it like this
"This isn't really asking me to list eval ideas. It's asking, with five real work days, which single mistake would you be most upset to find out about three weeks late, and can you catch that one first."
Why this works
Stops you from reciting a checklist of eval types instead of naming the one that actually earns the week.
4
Give the one decision, plainly
Say it like this
"I'd spend the week building one thing: a golden set of thirty to forty real transcripts, graded by someone who was actually in the room, scoring whether each action item has the right task, the right owner, and the right deadline. Not a broad quality eval across everything Notchbrief writes, just that one output, because it's the one nobody rereads before it gets acted on."
Why this works
This is the actual answer to the question, in one breath.
5
Prove it with the failure, cut to four sentences
Say it like this
"Here's what happens without that. Notchbrief's pilot ran fine for nine weeks with an associate reading every draft before it went to a client. Leadership decided to cut that read-through starting Monday, to keep the promise that action items land in your inbox before you're back at your desk. Nobody had a number for how often Notchbrief got the action items right on its own, only a feeling that it seemed mostly fine."
Why this works
Shows the real cost of skipping the check, not just the mechanism behind it.
6
Say what you would measure going forward
Say it like this
"After launch, I'd watch how often a client or teammate replies days later asking 'wait, who was supposed to do this,' tracked apart from every other kind of reply, since that's the closest thing to a complaint a missed action item ever generates."
Why this works
Shows you're thinking past the week, and names the one signal that stands in for a complaint nobody files.
7
Say what you'd leave alone
Say it like this
"I wouldn't spend week one grading Notchbrief's writing style or how it summarizes small talk. A stiff sentence gets reread and shrugged off. It's not the mistake that costs anyone a client."
Why this works
Shows judgment instead of trying to test everything evenly under a deadline that doesn't allow it.
8
Close on the decision, not the story
Say it like this
"So: with one week, build the golden set for action item correctness first, graded against real calls by someone who was there. That's the mistake nobody catches until it's too late to undo cheaply."
Why this works
Ending on the rule, not the anecdote, is what makes this sound like a method you'd actually reuse.
Let's learn
Notchbrief listens to a client call, and about ninety seconds after it ends, everyone on the call gets a summary and a list of who owes what by when.
Before Notchbrief, an associate at Fenshaw and Corbeau spent about twenty five minutes after every client call writing the notes and the action items up by hand, across roughly thirty five client meetings a week, team wide.
Knowledge spark: what's a golden set?
A small stack of real examples where a person has already worked out the correct answer by hand. You grade the tool against that stack instead of guessing whether its output looks right.
With Notchbrief running but an associate still reading every draft before it went out, the draft was ready in ninety seconds, and the read-through added back about eight minutes, catching a wrong name or a misassigned task in a couple of lines. That eight minutes felt like plenty of safety margin, so nobody built a real eval. Why would you, when a person reads it every single time.
We didn't need to catch every mistake. We needed to catch the one nobody was going to notice on their own.
Now leadership is cutting that read-through step, starting Monday, one week from today, to keep a sales promise that action items land in your inbox before you're back at your desk. The extra mistakes Notchbrief makes on its own are not really the problem. Most of them are typos and clunky sentences, and someone notices and fixes those the moment they open the email. The real problem is the mistake nobody is specifically looking for: an action item assigned to the wrong side of the table, or invented outright, sitting in a list that reads exactly as confident as the correct ones.
Day one check: how often each mistake type got caught without anyone being told to look
Self-caught rate, wording and summary slipsSelf-caught rate, action item mistakes
On a hand check of ten recent drafts against their real transcripts, a wording slip got noticed and mentally fixed 88 percent of the time. A wrongly owned or invented action item got noticed on its own only 9 percent of the time. Same tool, two very different kinds of miss.
The golden set is the one box that has to exist before anything downstream of it can happen, and it's buildable with what Beckett already has on day one.
At its worst, Notchbrief without a gate is worse than never building it at all. Two people each think the other one owns a task. Nobody does it. By the time anyone notices, three weeks have gone by, a deal has gone quiet, and nothing in the notes was ever flagged as unusual.
The choice that mattered
When Notchbrief first launched as an internal pilot, the team decided eval could wait until they had "enough real usage" to build a proper golden set. That made sense before launch, when there were only a handful of real transcripts sitting around. It stopped making sense the day leadership picked a Monday to remove the one thing standing between the model and the client.
What I'd leave alone: the summary's tone and phrasing don't get the same scrutiny in week one. A slightly stiff sentence costs a raised eyebrow, not a missed deliverable.
The lesson: a mistake with no complaint attached to it is not a small mistake. It's a mistake nobody's counting yet, and a week is exactly enough time to start counting the one that matters.
Now here is the same thing as a story
Read the short version above for the two minute answer. Read this when you want to feel why waiting for "enough usage" felt like the reasonable call at the time.
Before Notchbrief existed, Beckett Prentice was one of the associates at Fenshaw and Corbeau who wrote up call notes by hand, every evening, for years. Nobody's action item ever went to the wrong person on Beckett's watch, because Beckett had sat in the room and knew exactly who'd said they'd send what.
Notchbrief launched as an internal pilot, and the good months were good. A draft landed ninety seconds after a call ended instead of twenty five minutes later. An associate still read every one before it went to a client, caught the odd wrong name, fixed it, sent it. Nobody thought twice about skipping a formal eval. The reading step was the eval, in everyone's head.
Then the habit thinned, in three small beats. First, reading every line against the actual call became skimming the top summary and trusting the action item list, since it had been consistently fine for weeks. Then, associates rushing between back to back calls started forwarding drafts straight to clients without the full read, just this once, just today. Then, the daily standup stopped asking "did you check the action items specifically," and it quietly stopped being a step at all.
The trigger was small. A new associate Beckett was training asked a question nobody had a real answer for: "How would we even know if Notchbrief assigned a task to the wrong person, since nobody's going to complain about a task they don't know is theirs?" Beckett didn't have a good answer. Said he'd get back to her.
The read-through had caught every mistake anyone could see. It never asked what happens the week nobody's reading it anymore.
Before Beckett could build anything, leadership set the date: the read-through was being cut starting Monday, to keep the sales promise that action items land in your inbox before you're back at your desk. That left five work days. Beckett spent day one hand checking ten recent drafts against their real transcripts, the same check that produced the eighty eight versus nine percent gap. That told him where to point the rest of the week: not at wording, at ownership.
By Thursday, the full golden set was forty real transcripts, graded by Beckett and one associate who'd sat in each call. Recut by the kind of task, a pattern jumped out. Action items involving a hand-off, one side sending the other side data or a document, were misattributed a third of the time, thirty three percent, wrong owner or invented outright. Every other kind of action item sat at four percent. The blended number across all action items had looked fine. The hand-off slice alone did not.
Misattribution rate by action item type, on the finished forty transcript golden set
A blended score across every action item would have shown a comfortable low number. Recut by type, one category, hand-offs, carried nearly all of the risk.
The decision that opened the door went back to that first launch meeting. The team agreed eval could wait until Notchbrief had "real usage" behind it. Nobody wrote down what "real usage" actually meant, or set a date to revisit it, because at the time it felt like a sensible thing to defer, not a live risk.
Run the same week again with one change: Beckett doesn't wait for a full rebuild of the review step, and doesn't cut it entirely either. He keeps a short human check only on action items that involve a hand-off, roughly one in three of the week's drafts, and lets everything else go straight out. Review time drops from eight minutes on every one of the week's thirty five drafts to about ninety seconds on the third that actually carry the risk. On the following Monday, that narrow check catches four hand-off items that would have gone to the wrong side, before a single client ever sees one.
One design trusted a single person's eyes on every draft to speak for a mistake type nobody had actually measured. The other design measures the one category that matters and keeps a person only there.
What I'd tell myself, back when the new hire asked her question: the day someone asks how you'd know, that's the day to start counting, not the day to promise you'll get back to them.
ORDER, when the week is the whole budget
Not a checklist for a planning meeting. Five questions that build toward the one that actually decides what gets built first: what's hardest to undo.
OOutcome. What is every candidate eval actually competing to move?
Not "does Notchbrief sound smart." Whether its output can be sent straight to a client with nobody reading it first and still be trusted enough to act on.
Without naming this, ranking eval ideas is just a matter of taste.
RReversibility. Which mistake is hardest to undo?
A wrong sentence in the summary gets reread and fixed in a minute. A wrongly assigned or invented action item doesn't get noticed until the task nobody did surfaces weeks later, and by then the fix isn't a correction, it's an apology.
This is the step that decides everything else. If a mistake gets caught and undone for free, it doesn't need the week. If it doesn't, it's the whole reason the week exists.
One of these gets caught by a stranger rereading it. The other one never gets reread at all.
DDependency. What has to exist before anything can be graded?
Real transcripts and two people who were actually on the calls, before a single score can be computed. Both are buildable inside the week itself, unblocked by anything outside it, which is exactly why they're the right thing to spend day one on.
A plan that starts with "wait for more usage data" isn't blocked by reality. It's a choice dressed up as one.
EEvidence. What can you learn cheaply before committing the rest of the week?
Day one: hand check ten recent drafts against their real transcripts before building the full golden set. That's what turned up the eighty eight versus nine percent self-caught gap, and it's what pointed the other four days at ownership instead of wording.
Spending a whole week grading blind, without a cheap first look, is how you end up polishing the wrong output.
RRank. State the order, and defend the top pick.
Golden set on action item correctness first. Inside it, invented tasks and hand-off misattribution scored as their own category, not folded into a general "wrong owner" number. Summary tone and phrasing waits for week two.
The top pick wins because it's the output people act on unread. Everything else on the list gets a second look for free; this one doesn't.
Three things worth stating directly, since this is where the real judgment sits. The rejected alternative was a broad, automated eval scoring every kind of thing Notchbrief writes, tone, structure, and action items alike, in the same week. It lost because a five day budget spread evenly across everything ends up an inch deep on the one output where being wrong actually costs someone a client. The AI specific failure worth naming by name is the invented action item: a task that reads exactly as confident as a real one but was never agreed to on the call, which is a different and worse failure than a task just sent to the wrong person. The guardrail is scoring invented tasks as their own tagged category inside the golden set, not blended into a general accuracy number. And the trade being accepted on purpose is real: keeping a short human check only on hand-off type action items, instead of cutting the human step to zero or keeping it on every draft, costs real minutes on roughly a third of the week's output, traded against a category that had shown thirty three percent misattribution with nobody watching for it.
The five, in one line each: O: whether the output can be trusted enough to act on unread. R: a wrong sentence gets reread for free. A wrong action item doesn't. D: real transcripts plus two people who were in the room, buildable inside the week itself. E: a cheap ten-draft hand check on day one points the rest of the week at the real risk. R: action item correctness first, hand-off tasks flagged on their own, tone waits.
Same five letters, a vet clinic instead of a client call
Not every silent mistake sits in a client inbox. Sometimes it sits on a kitchen counter, in a bowl labeled twice a day instead of once.
Coastwell Veterinary Partners runs Scriblatch, a tool that listens to a vet visit and, once the appointment ends, texts the pet owner a summary and a list of home care tasks, medication, dose, how often, for how long. Griffen Askew runs eval on it, also with one week before the clinic's front desk stops reading every text before it sends.
Griffen's day one hand check, ten recent visits against the vet tech's own written notes, found something close to Beckett's split. A wrong clinic detail or a stiffly worded line got caught by an owner calling in to ask, almost every time. A wrong medication dose or frequency didn't get caught by anyone, not until the follow up visit, sometimes two weeks later.
The decision Griffen would rank first
A golden set of real visit transcripts, thirty to forty of them, graded against the vet tech's own dosing notes, scoring only drug, dose, frequency, and duration. Not tone, not the small talk summary, because a wrong dose doesn't get reread and corrected. It gets given.
Same method, different shape: a wrongly assigned meeting task and a wrongly stated medication dose look nothing alike, but both are a model output that gets acted on by someone who has no reason to double check it, before anyone outside the company would ever know it was wrong.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the one line: find the mistake that doesn't get reread, then spend the week building the eval that catches that one first.
Cost: there's no budget this quarter for a full labeling team. Grade fifteen transcripts instead of forty, with the two people you already have. The sample shrinks; the category you're protecting doesn't change.
The model got better, for real: say the newest version is measurably more accurate on average. That's exactly when this matters most, since a more accurate model still makes its rarest mistakes with full confidence, and a rare, confident mistake is the hardest kind for a person to catch by eye.
Where people run it wrong.
They build one broad eval across every output instead of ranking by which mistake actually can't be undone.
They grade against a clean, written up transcript instead of the messy real recording, and miss the exact failure that only shows up in overlapping, real speech.
They wait for "enough real usage" before starting the golden set, instead of using the handful of real transcripts already sitting in a folder.
How to use it live. Say the ranking rule out loud before you say anything else: "rank by what's hardest to undo." That buys a beat of thinking time, and it turns the rest of the answer into naming the one mistake that fits that description.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
ORDER: rank by what's hardest to undo. Built for prioritization-under-constraint questions, when there isn't time or people to do everything at once.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Beckett Prentice, eval lead on Notchbrief at Fenshaw and Corbeau, a consulting firm. Wrote call notes by hand for years before the tool existed.
3 · THE HABIT
What did the team stop doing because it worked?
Tap to flip
ANSWER
They stopped fully rereading the action item list against the call, since an associate's read-through step had made a full recheck feel unnecessary.
4 · THE HARD CALL
What's the reversibility comparison this whole answer turns on?
Tap to flip
ANSWER
A wording mistake gets reread and fixed the same day, for free. A wrongly owned or invented action item goes out unread and doesn't surface until the task nobody did shows up weeks later.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Deciding, at launch, that eval could wait until they had "enough real usage," instead of building a golden set from the real transcripts already on hand.
6 · THE NUMBER
Fill in the blank: on the day one check, wording slips were self-caught ___ percent of the time, versus only ___ percent for action item mistakes.
Tap to flip
ANSWER
88 percent versus 9 percent. The finished golden set later found hand-off type action items misattributed 33 percent of the time, against 4 percent for every other type.
7 · THE REPLAY
Same one week, new design, what changes?
Tap to flip
ANSWER
Beckett keeps a short human check only on hand-off type action items, about one in three drafts. Review time drops from 8 minutes on every draft to about 90 seconds on the ones that need it, and it catches 4 misassigned hand-off tasks before Monday's launch.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the one eval it ranks first?
Tap to flip
ANSWER
Scriblatch, a vet visit tool at Coastwell Veterinary Partners. Griffen Askew ranks a golden set on medication instruction correctness, drug, dose, frequency, duration, first.
Check yourself Score: 0 / 0
Multiple choice
1. With only one week, which eval gives you the most protection against the costliest mistake?
A. A broad automated eval scoring writing quality across every Notchbrief output.
B. A small golden set of real transcripts scoring whether each action item has the right task, owner, and deadline.
C. A survey asking associates how much they trust Notchbrief.
D. A speed test measuring how fast Notchbrief responds after a call ends.
Show hint
Think about which mistake type gets caught for free, and which one doesn't.
Show answer
B. Action items get acted on without a second look. A wrong one is the mistake nobody's watching for, so it's the one worth the week.
True or false
2. True or false: because Notchbrief's overall accuracy already looked fine to the team, the action item mistakes weren't worth testing for in week one.
True
False
Show hint
Check the bar chart in Section 1. Look at the self-caught rate for action item mistakes on its own.
Show answer
False. A general "seems fine" feeling hid the fact that action item mistakes were self-caught only 9 percent of the time, against 88 percent for wording slips. A comfortable overall impression can sit right on top of a badly unwatched category.
Fill in the blank
3. Fill in the blank: on the finished golden set, action items involving a hand-off were misattributed ___ percent of the time, against ___ percent for every other type.
Show hint
The number is in Section 2's story and repeated on flashcard 6.
Show answer
33 percent versus 4 percent. A blended score across all action items would have looked comfortable. Recut by type, one category carried nearly all the risk, which is exactly what the E step in ORDER is built to find cheaply.
Short answer, name the rejected alternative
4. What did Fenshaw and Corbeau's team decide about eval when Notchbrief first launched, and why did that decision stop making sense the week leadership set a cutover date?
Show hint
Look at the block-key box titled "The choice that mattered" in Section 1.
Show answer
Model answer: They decided eval could wait until they had "enough real usage" to build a proper golden set. That made sense before launch, with only a handful of real transcripts. It stopped making sense once leadership scheduled the removal of the human read-through, the actual backstop, with no eval built to replace it.
Multiple choice
5. The golden set found hand-off type action items misattributed 33 percent of the time while every other type sat at 4 percent. What does that combination tell you?
A. The tool is working fine overall, since most action item types stayed low.
B. A blended accuracy number would have hidden a category carrying nearly all the real risk, which needs its own check, not an average.
C. Every action item should get the same heavier scrutiny, evenly, going forward.
D. A low blended rate proves no single category needs special handling.
Show hint
Look at what recutting by task type showed that the blended number hid completely.
Show answer
B. A steady blended number hid a category specific spike. Spreading scrutiny evenly across every type would have wasted most of the week's review time on categories that barely needed it.
Short answer, apply it yourself
6. Pick an AI product you use yourself where a wrong output might not get checked by anyone unless it resurfaces later. With only a week, what's the one eval you'd build first for it?
Show hint
Think of a product where the output gets acted on, not just read, a delivery app, a scheduling tool, an auto-reply.
Show answer
Model answer: A grocery delivery app's substitution list, where the app picks a replacement item when something's out of stock. Nobody rereads the substitution before it's packed. I'd build a golden set of real orders scoring whether each substitution respects the stated allergy or dietary flag on the account, since that's the one substitution mistake nobody catches until it's already in the bag.
Before you close the answer
Why this works
Tests whether you protect the mistake nobody complains about, or spend a scarce week polishing the part people already notice and fix themselves. Most candidates list several eval ideas and never actually rank them.
Follow-up traps
"Why not build a broader eval so you're covered on everything?" Response: a week graded broadly ends up an inch deep on the one output where being wrong costs a client. Better to fully cover the highest stakes slice than partly cover all of them.
"Isn't forty transcripts too small a sample to trust?" Response: yes, on its own. That's why it's paired with a lightweight human check on the risky category rather than treated as proof the tool's safe everywhere. It gets sized up in week two once real production volume comes in.
If pressed
The "hand-off" category wasn't a gut call by one person. Before grading, Beckett and the associate agreed on a written rule: any action item whose sentence implies a document, number, or file crossing from one side of the call to the other counts as a hand-off. Writing that rule down before scoring is what made the 33 percent a real finding instead of one grader's impression.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.