Describe the sampling strategy for building an eval set from production traffic.
- Pull a fixed number from every category, not a random sample weighted by traffic.Why: a random pull gives you almost nothing from a category that's one percent of your traffic, even when that one percent is the one that can get a client sued.
- Set a real floor per category, big enough to catch a real pattern, not just "a few examples."Why: three or four examples can't tell a one-in-ten failure rate from noise; you need enough for the pattern to show itself.
- Double the floor for categories where a miss is a legal or brand risk, not just an off note.Why: a caption that guesses wrong about a lunch special is a shrug; a caption that makes an unproven health claim is a lawsuit sitting on a client's desk.
- Size the total against how long a person can actually read it, in one sitting or one week.Why: a number that sounds rigorous but nobody can review before the next release isn't a real eval set, it's a slide.
- Know whether the category count or the per-category floor is doing the heavy lifting in your total.Why: cut the wrong one and you either blow the review budget or thin out the exact categories you added the floor to protect.
- Re-check the category list whenever a new kind of post shows real volume in production.Why: a taxonomy that stops updating quietly turns back into the same blind spot you just fixed.
How to answer this, stage by stage
Nobody is grading whether you land on exactly 640. They're grading whether the floor per category is big enough to mean something, whether the risky categories get more than an equal share, and whether the total survives a real review day. Eight moves get you there.
Let's learn
What happens when a sample size sounds big, and the one category that matters most only shows up six times?
Hashloom looks at a photo and a business's own notes, and writes the caption and hashtags for a social post, so the business doesn't have to think of one.
Before the team built any real eval set, they tested Hashloom on about 40 examples someone had written by hand, roughly a day's work, mostly picked because they were interesting, not because they were common.
Now Hashloom logs every post it actually generates in production, thousands a day. Pulling from that felt like the fix: a random sample of 500 real posts, weighted by how much traffic each kind of post actually gets.
Five hundred sounds like a lot more than forty. But that's not the real improvement, and it isn't the real problem either.
Here's the turn. A random pull weighted by volume gives you a sample that looks exactly like your traffic. Food posts are common, so you get plenty of food examples. Health and wellness claims are rare, about one post in seventy, so out of 500 you get maybe six or seven. Alcohol posts are rarer still. You end up with a big, thorough-looking sample that has almost nothing to say about the two categories where a wrong caption is a legal problem instead of an embarrassing one.
At its worst, that eval set passes every release with a strong score, and a false health claim ships anyway, because there were never enough health-category examples in the set for the pattern to even show up.
The choice I would take back. The team sized the eval set as one random pull, 500 posts, weighted by real traffic volume, because that felt more honest than hand-picking examples. It was fine while every category carried about the same risk. It stopped being fine the day one caption used the word "proven."
What I would leave alone. The random pull is still fine for something like caption tone, where every category behaves about the same and a miss just looks a little off. There's no need to stratify sampling for a low-stakes read like "does this sound like the brand," only for the categories where a miss actually costs someone something.
The lesson. A sample size that sounds thorough is not the same as a sample that has actually seen the thing you're worried about. If the categories with the least traffic are also the ones with the most risk, a sample built to match your traffic will always be the last one to notice.
Now here is the same thing as a story
Skip this if you already believe a random sample and a stratified one aren't the same thing. Read on if you want to feel why they aren't.
Fumiko Arai has run product for Hashloom's caption engine for two years, and she can tell you which categories of small businesses post the most without looking anything up: food and drink first, fashion close behind, fitness a clear third.
She built the first real eval set herself, the week Hashloom crossed ten thousand captions a day. Five hundred posts, pulled at random from that week's real traffic, weighted the way the traffic actually looked. It felt like the honest choice. Nobody had cherry-picked a single example.
For months, that felt like enough. Every release, she'd run the 500 against the new model, read through the flagged ones over a Tuesday afternoon, maybe two hours of work, and ship. The pass rate held in the high nineties. The first couple of times, she spot-checked the health and alcohol posts in that set specifically, out of habit more than worry. There were only ever six or seven of them in the whole sample. She'd read all six, they'd be fine, and she'd move on.
So she stopped opening those six on purpose. Then she stopped thinking about the split by category at all. Five hundred examples, one pass rate, one number to watch.
Then came a Thursday, a message from a client's own legal team, not a support ticket. A supplement brand had posted with a Hashloom caption reading, more or less, "Clinically proven to melt stubborn belly fat in one week." Nobody at Hashloom had written that sentence. The model had, on its own, from a product photo and three lines of the business's notes about what the supplement did. The brand's lawyer caught it before it went live and asked, reasonably, how a tool that generates this many captions a day let that one out.
Fumiko pulled the eval set that had passed the release the caption shipped under. Five hundred posts. Six were from the health and wellness category. None of the six had anything close to an unproven claim in them, because six examples from a category that generates hundreds of captions a day was never enough to show what that category's failures actually look like.
She went back to the eval and asked it the question the 500-post pull had never been built to answer. She pulled every health and wellness caption from three weeks of real logs, about 340 of them, and read a stratified slice of 80. Eleven of the 80, about 14 percent, made a claim no honest business could stand behind. Not because the model had gotten worse. Because nobody had ever looked at that category closely enough to know the problem was already there.
She rebuilt the eval that week: a floor of 40 examples from every one of Hashloom's fourteen content categories, doubled to 80 for health and wellness claims and for alcohol and age-restricted products, 640 examples in total, about 16 hours for one reviewer to get through. Run against the model as it stood, the health and wellness slice scored 86 percent clean, well under the bar she set for it. A fix requiring the model to flag any sentence with a medical or outcome claim for a lighter, hedged rewrite took a week. Rerun, the slice cleared 98 percent. Over the following month, Hashloom generated just over 6,000 health and wellness captions with zero flagged claims reaching a client unreviewed.
The thing I'd tell myself, the week I built that first 500-post sample: a number that looks thorough because it's big is not the same as a number that's actually seen the thing you're worried about. Ours had five hundred examples, and had barely looked at the six that mattered.
BOUND, run against a category with six examples in it
This is a sizing question about how many examples to pull and from where, not a person's trust flipping between two settings, so BOUND fits and FLIPS doesn't.
B, break it down. Sample size equals the number of content categories, times a minimum floor per category, plus extra for the categories already known to carry more risk than the rest. A typical Hashloom taxonomy has 14 categories.
O, own the numbers. A floor of 40 examples per category, since under 40 a real one-in-ten failure rate can't be told apart from noise. 14 times 40 is 560. Two categories, health and wellness claims and alcohol and age-restricted products, get that floor doubled to 80, adding 80 more. Total: 640.
U, use a range. With a coarser taxonomy, 8 broad categories at a flat 40, the total drops to 320. Splitting by category and by platform too, 14 categories across 4 platforms, pushes it past 2,000, closer to 2,560 once the flagged categories stay doubled inside every platform. Start at 640, the plain category split, and only add platform as its own axis where the model's behavior genuinely differs by platform.
N, nail the sanity check. 640 examples at about 90 seconds each to check the caption, the hashtags, and any unproven claim is 16 hours. Two working days for one reviewer, or a single day split across two or three. A real pass someone can finish before the next release, not a number that only sounds rigorous.
D, direction. The category count moves the total more than the floor does. Growing the taxonomy from 14 to 20 categories adds about 240 examples. Raising the floor from 40 to 45 only adds about 80. If the total needs to shrink, merge overlapping categories before thinning the floor, since a thin floor is exactly what let the health claims category hide the first time.
And if you want to be sure it really works, try it somewhere else
A vet-clinic tool drafts the visit note summary from what a vet says out loud during an appointment, turning spoken observations into a written record with medications and doses.
B, break it down. A typical taxonomy is species times visit type: 3 species buckets, dog, cat, exotic and other, times 5 visit types, wellness, sick, emergency, surgery follow-up, dental, for 15 strata.
O, own the numbers. A floor of 25 examples per stratum, 15 times 25 is 375. The exotic-species-plus-emergency stratum gets the floor doubled to 50, adding 25 more. Total: 400.
U, use a range. Splitting species more finely, say 6 species groups instead of 3, doubles the stratum count and pushes the total toward 750. Fewer visit types, just wellness, sick, and emergency, brings it down closer to 225. Start with the 15-stratum version and widen the species list only where a specific exotic species shows real visit volume.
N, nail the sanity check. Checking a visit summary for dosage accuracy takes longer than checking a caption, closer to 3 minutes each. 400 examples is 20 hours, about two and a half days for one reviewer, or split across a vet and a technician in one day.
D, direction. Here the category count isn't the real lever. A geriatric cat on three medications during a routine wellness visit carries the same dosage risk as an exotic animal in an emergency visit, and stratifying only by species and visit type misses that entirely. The floor that matters most is a separate flag for multiple-medication cases, not a finer split of species or visit type.
Swap the trigger and it still runs.
Speed: an interviewer caps the answer at 90 seconds. Skip straight to the split: 40 per category, doubled to 80 for health claims and alcohol, 640 total, 16 hours to review. The build-up backs it up if they ask.
Cost: instead of asking what the sample size should be, a manager caps review time at one day, 8 hours. Work backward: 8 hours at 90 seconds a piece is about 320 examples. Merge overlapping categories to fit that ceiling before thinning the floor on health claims or alcohol.
The model got better: a new version almost never makes a false health claim anymore in general use. The floor doesn't drop on its own. Rerun the health and wellness slice specifically, since a clean overall number says nothing about the category that used to fail.
Where people run it wrong.
They size the sample by what sounds like enough, 500, 1,000, instead of a floor times a category count.
They set the floor once at launch and never revisit it when a new category, a new regulated vertical, starts showing real volume.
They stratify by whatever's easy to tag, platform, post length, instead of the thing actually tied to risk, the content category or the claim type.
How to use it live. Say the equation before any number: "sample size is category count times a per-category floor, doubled wherever a miss actually costs something." That buys the time to work out the real floor instead of guessing a round number that sounds thorough.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Writing an eval spec
- #1 What is an eval spec and who is its audience?
- #2 List the components of a complete eval spec.
- #3 How do you define a task-level success criterion for a summarization feature?
- #4 Write a scoring rubric for the quality of a generated customer support reply.
- #5 Describe the difference between an eval spec and a test plan.
- #6 How many examples belong in a first eval set and how do you choose them?