Artifact critiqueIntermediateQuality, Cost & Token Economics / Eval design for product teams / #6

Describe a rubric that a non-technical reviewer could apply consistently.

A five point "does this sound appetizing" scale asks a stranger with no training to trust their own gut. A rubric built to last asks one thing only: is this claim true about this exact dish.

The direct answer
Build the rubric out of yes or no items that check one fact about this exact output against its own source material, for example "does this description name an ingredient that isn't on the dish's ticket." Never an item that asks the reviewer to rate how good, appealing, or appetizing something sounds, because two honest people will always disagree about a feeling and only one of them can be checked against anything real. On Garnish's first rubric, a five point "how appetizing" scale, two restaurant owners agreed on the same batch of descriptions only 61 percent of the time and caught real planted errors only 46 percent of the time. Swapping every item to a checkable fact pushed both numbers past 90 percent, with the same reviewers, on the same clock.
Do this, in order
  1. Write every rubric item as a yes or no fact you can check against the dish's own ticket, never a 1 to 5 feeling.Why: a feeling has no ticket to check it against, so two honest reviewers land in two different places every time.
  2. Build a small golden set with planted errors before you ship any version of the rubric.Why: you can't know a rubric works, you can only measure it. A rubric that "feels thorough" with no test set is a guess wearing a lab coat.
  3. Track cross reviewer agreement as its own number, separate from how many descriptions get approved.Why: two reviewers can approve the same batch at a healthy rate and still be catching completely different things, or nothing at all.
  4. Keep the rubric to four or five items on one screen, no scrolling.Why: an owner closing up at 11pm will skim past item nine, and a skimmed item catches nothing.
  5. Route anything needing training the reviewer doesn't have, legal wording, allergen disclosure rules, off the rubric to a specialist queue.Why: asking an untrained person to judge a legal claim doesn't protect anyone, it just hides the risk behind a box someone ticked without knowing what they were ticking.
  6. Leave spelling and grammar off the rubric entirely.Why: an automated proofreader already catches that, and every item you add for something handled elsewhere steals attention from the item that actually matters.

How to answer this, stage by stage

Nobody's grading whether you can describe a review process. They're grading whether you can name one design decision precise enough that a stranger could argue with it.

1
Ground it in one dish, one reviewer, before you say the word rubric
Say it like this
"Let's ground this in Garnish. It takes a short ingredient list, say grilled salmon, lemon butter, asparagus, and writes a menu description ready to publish. Alva Faber owns Birchwood Tavern, one restaurant, and she reviews every description herself before it goes on her online menu."
Why this works
An abstract "design a rubric" question turns into a lecture about quality processes fast. One owner, one dish, makes it a real decision instead.
2
Name what review looks like today, and why it can't just be inherited
Say it like this
"Before Garnish, Alva wrote every description by hand at closing time, about twelve minutes a dish. When Garnish took over the writing, nobody actually designed the reviewing. It just inherited whatever felt natural, which was rating how the sentence sounded."
Why this works
Naming the situation stops you designing a rubric in a vacuum. It has to fit a real closing shift, not an imagined one.
3
Say the one habit the rubric has to build
Say it like this
"The habit I want isn't 'read carefully.' It's narrower. I want Alva answering three or four fixed yes or no questions about this exact dish, using the ticket in front of her, instead of asking herself whether the paragraph sounds nice."
Why this works
A vague goal like "review it well" produces a vague rubric. A named habit produces named items.
4
Give the anchor, the actual design decision
Say it like this
"Here's the anchor. Every item on the rubric has to be a fact you can check against the dish's own ticket, never a quality judgment. 'Does this description name an ingredient that isn't on the ticket, yes or no' is a rubric item. 'Does this sound appetizing, one to five' is not. That's a taste test wearing a rubric's clothes."
Why this works
This is the actual answer to the question. Everything else in the walkthrough is defending it.
5
Show what breaks the first time an item stays vague
Say it like this
"Garnish's first rubric had exactly one item: rate how appetizing this sounds, one to five. A description for Birchwood's summer salad said it was finished with a scatter of toasted hazelnuts. There were no hazelnuts on the ticket. Alva liked how it read and gave it a five. A guest ordered it for the hazelnuts, got a plain salad, and left a review calling the menu made up."
Why this works
One vague item is enough to let a real error through, because a reviewer can only check what the item actually asks them to check.
6
Say what you leave off, on purpose
Say it like this
"I'd keep two things off this rubric completely. Anything needing a legal read, like health or allergen wording, goes straight to Garnish's compliance queue, not to Alva. And spelling stays with the automated proofreader that already runs on every draft. Alva's items are the ones only she can answer, because only she knows what's actually on the ticket."
Why this works
A rubric that asks a non-technical reviewer to judge something they were never trained to judge doesn't protect anyone. It just moves the risk somewhere invisible.
7
Close on the measured proof, not a promise
Say it like this
"So: after the switch, the same two reviewers who agreed on a batch of descriptions 61 percent of the time under the old scale agreed 94 percent of the time under the checklist, and caught 22 of 24 planted errors instead of 11. Check facts against the ticket, not feelings against a scale."
Why this works
Closing on a measured number, not "it worked well," is what makes this sound like a shipped design instead of a nice idea.

Let's learn

Every night Birchwood Tavern closed, Alva Faber used to sit at the bar with a legal pad and cross out any description that oversold the food, about twelve minutes a dish.

Garnish is a tool that takes a short ingredient list, say grilled salmon, lemon butter, asparagus, and writes a description ready for the menu board. Before it, Alva either wrote every description herself, or paid a freelance copywriter roughly 35 dollars a plate for the ones she didn't have the evening for.

Now Garnish drafts a description in about four seconds. Alva's job doesn't disappear. It changes shape, from writing to reviewing, and nobody at Garnish Labs had actually designed what reviewing was supposed to look like.

The turn isn't that Garnish sometimes gets a detail wrong. Every model does. The turn is that the first rubric Garnish shipped couldn't tell you when that happened, because it never asked Alva to check anything. It asked her to rate a feeling.

That first rubric was one line: rate how appetizing this description sounds, one to five. Garnish Labs tested it against a golden set, 150 already written descriptions a food copywriter had graded by hand, 24 of which had one ingredient planted in the text that wasn't really on that dish's ticket. Two restaurant owners scored the same 150 independently.

Knowledge spark: what's a golden set? A batch of outputs someone has already graded by hand, on purpose, and kept aside just for testing. Some of the batch has known problems baked in, like a planted wrong ingredient, so you can measure whether a rubric actually catches what it's supposed to catch, instead of guessing.

They agreed on which ones to approve only 61 percent of the time. Of the 24 planted errors, they caught 11.

We didn't lose thirteen missed errors. We lost the one thing a rubric is supposed to do: make two honest people land in the same place.

At its worst, that gap cost more than one wrong plate. A description for Birchwood's summer salad, mixed greens, goat cheese, dried cherries, sherry vinaigrette, said it was finished with a scatter of toasted hazelnuts. There were no hazelnuts on the ticket, not close. Garnish had simply made the line up because it read well. Alva rated it a five. A regular guest ordered the salad specifically for the hazelnuts, got a plain one, and left a public review calling the online menu made up. Alva pulled all 40 descriptions for that summer menu offline for a week and rechecked every one by hand, which is exactly the twelve minutes a dish she'd hired Garnish to save her.

The choice that mattered Months earlier, when Garnish Labs built that first review screen, the team had two options on the table: a plain approve or reject toggle, or a fuller one to five "how appetizing" scale that product wanted, because it could double as an early signal for a future feature that would recommend the most enticing item on a menu. The richer scale felt like it was doing more work. Nobody separately asked whether it also had to catch a description that was simply wrong.
Old scale vs. checklist: agreement and error catch rate
100% 50% 0% 61% 94% 46% 92% Reviewer agreement Planted-error catch rate
Old scale, "how appetizing, 1 to 5"Checklist, checkable facts against the ticket
Same 150-description golden set, same two reviewers, both times. Only the wording of the rubric changed.

What I'd leave alone: Garnish's spelling and grammar. An automated proofreader already checks every draft before it reaches Alva, and in six months it hasn't missed a typo worth mentioning. Adding a spelling item to her rubric would just be one more line for her to skim past on a night she's already tired.

The lesson: a rubric can be filled in carefully, by someone who genuinely cares about getting it right, and still teach nothing, because a vague item doesn't test the output. It tests the reviewer's own taste, and calls that a check.

Now here is the same thing as a story

Read the long version below when you want to feel why a five point scale can be filled in honestly and still be useless, not just be told that it is.

Aslak Kessing runs quality at Garnish Labs. He built the review process almost by himself in the company's first year, and he still keeps a running note of every dish description reviewers have ever flagged, going back to launch.

For the first several months after the one to five scale shipped, Monday mornings were easy for him. He'd pull the weekend's approval numbers, see them holding steady around 90 percent approved, and move on to the next fire. Restaurant owners liked the scale. It felt quick, and it felt like their own judgment mattered.

At first, owners read the whole description before scoring it. Then most started reading the first line and letting it decide the score. By month five, a handful of owners, busy ones, mostly single location owners running the review themselves at the end of a long shift, were giving out fours and fives on autopilot, the way you nod along in a meeting you stopped following twenty minutes ago.

The trigger for Aslak wasn't a complaint. It was a new hire on his team doing a routine spot check, comparing how three different restaurants scored the exact same stock phrasing, the same sentence shape Garnish tends to reuse for salads. One owner gave it a two. Another gave the same structure, wrong ingredient and all, a five. The new hire asked Aslak why. He didn't have an answer.

So he pulled the golden set. 150 descriptions, 24 with a planted ingredient that wasn't on the real ticket. Two owners, scoring blind, agreed on approve or reject 61 percent of the time. They caught 11 of the 24 fakes. Not because Garnish was making things up constantly, it wasn't, but because the scale never once asked anyone to check.

The real cost wasn't the 11 descriptions that slipped through. It was every owner on the platform quietly learning that the review step was a formality, since it had never once caught anything that mattered to them personally.

It was never really about the 61 percent. Owners didn't have a number in their heads either. They had a feeling: this line reads fine, or it doesn't. The scale asked for the feeling directly, so that's exactly what it got back, nothing more.

The decision Aslak would take back happened in a room eight months before any of this. Someone on product wanted the plain approve or reject toggle Aslak's own gut preferred. Someone else argued for the fuller scale, because a number from one to five could later train a feature that recommended the most enticing dish on a menu to diners. The room picked the richer option. Nobody in that meeting asked whether the richer option also had to catch a lie.

Run the same spot check again with the checklist rubric in place. Same stock salad phrasing, same fake hazelnut line, three restaurants. Item one asks: does this name an ingredient not on the ticket. All three owners answer no, correctly, in under thirty seconds each. Across the full golden set, agreement climbs to 94 percent and the catch rate to 22 of 24, in the same four minutes a batch the old scale used to take.

One design asked owners what they felt about a sentence. The other asked them one fact about a plate. Only one of those can be checked.

What I'd tell my past self, sitting in that first meeting: the fuller scale wasn't wrong because it was ambitious. It was wrong because ambition and accuracy were never actually the same axis, and we built the whole review step as if they were.

Five letters, one plate

SPARK, run start to finish on the same rubric, so you can see where each letter actually earns its place.

SSituation. Who is this person, and how does the job get done today, without you?
Alva Faber, one restaurant, reviewing every description herself at closing time. Before any rubric existed at all, review was just "does this sentence sound right to me," inherited from nothing more than what felt natural.
Name the real shift this rubric has to survive, or the design floats free of the actual job.
Hand sketched flow diagram titled before Garnish, Alva's closing routine. Four connected boxes in sequence: close the floor, pull the ticket, write it by hand, chalk it on the board. The third box, write it by hand, is emphasized in mustard.
Twelve minutes a dish, every closing shift, before any AI touched the menu.
PPayoff. What habit do you want this to build?
Not "review carefully." Specifically: answer three or four fixed yes or no facts against the ticket in front of you, instead of asking whether the paragraph sounds nice.
A named habit produces named rubric items. A vague goal produces a vague rubric, every time.
AAnchor. The one design decision everything else hangs on.
Every item is a checkable fact about this exact dish, never a quality judgment. "Does it name an ingredient not on the ticket, yes or no" is the anchor line. "Rate how appetizing this sounds" is not a rubric item at all.
This is the actual design decision. If it doesn't visibly survive the next letter, it's not an anchor, it's a slogan.
Hand sketched comparison titled two ways to phrase one rubric item. Left panel a question mark box labeled vague item, caption does this sound appetizing 1 to 5. Right panel a document icon labeled checkable item, caption does it name an ingredient not on the ticket yes or no.
Only one of these two sentences can be checked against anything real.
RRisk. What breaks the first time you're wrong?
One vague item, "how appetizing does this sound," let a made up hazelnut garnish through at a five out of five. Two reviewers scoring the same output landed in two different places, because the item let each of them bring their own taste to it.
The anchor has to be built to survive exactly this. A rubric that only works when the model behaves is not a design, it's a hope.
Hand sketched comparison titled the day the model invents a garnish. Left panel a gauge icon labeled vague scale, caption reads nicely scored a 5 hazelnuts published none on the ticket. Right panel a document icon labeled checklist item, caption item 1 flags the ingredient mismatch before it ever posts.
The anchor is only worth anything if it still catches this exact failure.
KKeep out. What do you deliberately not build?
Anything needing training Alva doesn't have, a legal read on a health claim, an allergen disclosure format, stays off her rubric and goes to Garnish Labs' compliance queue instead. Spelling and grammar stay with the automated proofreader.
A rubric that asks an untrained reviewer to judge something they can't safely judge doesn't add safety. It just hides the risk behind a checkbox somebody ticked blind.
Hand sketched numbered list titled what stays off Alva's rubric. Item 1 health or allergen wording routed to legal review. Item 2 allergen disclosure format routed to compliance queue. Item 3 spelling and grammar handled by the automated proofreader.
Three things Alva never has to rule on, and three places they go instead.

Three things worth saying directly, since this is where the real judgment sits. The alternative the team rejected, twice, was a plain approve or reject toggle, which product turned down early on because it wanted a number that could later train a "most enticing dish" recommender. It lost, eventually, because a number nobody can check is not more informative than a toggle nobody can check, it just feels more sophisticated while doing the same job worse. The AI specific failure mode worth naming by name is the model inventing a garnish or ingredient that reads naturally but isn't in the source list, plain hallucination, dressed up as a nice sentence about hazelnuts. The guardrail is the ticket check item itself, backed by a golden set with planted errors so the team can measure catch rate instead of assuming it. The bar was never zero invented ingredients forever, a model answering from a short ticket will occasionally embellish, the bar is a checklist that catches it before a guest orders off it, calibrated to clear a 90 percent catch rate on the golden set before any new rubric version ships. The trade-off worth naming too: keeping legal and allergen wording off Alva's rubric and routing it to a specialist queue costs a day or two of turnaround on those specific items, slower than letting her tick a box herself, accepted on purpose because a wrong guess there costs far more than a slow one.

And if you want to be sure it really works, try it somewhere else

Same five letters, an animal shelter instead of a restaurant, nothing about menus anywhere in sight.

Kindly is a tool that drafts an adoption bio for a shelter animal from a short intake sheet: species, age, temperament notes, medical flags. Auberon Alderete manages a mid-size shelter and reviews every bio himself before it posts.

S, situation: before Kindly, a volunteer wrote each bio by hand from the intake sheet, about eight minutes an animal, on top of everything else a shelter shift already asks of someone.

P, payoff: the habit worth building isn't "does this bio feel true to the animal." It's answering a fixed set of yes or no facts against the intake sheet itself.

A, anchor: every item checks a named trait against the sheet, for example "does this bio claim a trait, like good with cats, that isn't marked on the intake sheet, yes or no." Never a feeling about whether the bio sounds right.

R, risk: the shelter's old rubric had one item, "does this bio feel true to the animal, yes or no," borrowed from a general customer service template because it sounded reasonable. Two volunteers gave opposite answers on the same bio for a dog whose intake sheet said untested with cats, because "feels true" was really asking for a stranger's ten minute impression, not a fact.

K, keep out: whether a temperament note is severe enough to flag a placement as high risk stays off the volunteer rubric entirely and goes to a vet tech or behaviorist. Volunteers check what the sheet says. They don't diagnose what it means.

The decision Auberon would take back Copying that first "does this feel true to the animal" item from a general training template instead of writing one from the shelter's own intake sheet. It sounded reasonable at the time. It let two volunteers land on opposite answers for the same bio, because a feeling about a dog they'd both only just met was never something either of them could actually check.
Hand sketched decision tree titled writing one item for a shelter bio rubric. Root question how do you phrase this rubric item. Left branch asks if the bio feels true leads to two volunteers two answers. Right branch asks if a trait is on the intake sheet leads to checkable same answer every time.
Same anchor, a different ticket. The sheet plays the part the ingredient list played at Garnish.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the anchor, every item is a checkable fact against the source data, never a feeling, and give the one number, agreement went from 61 to 94 percent.
Cost: there's no budget this quarter to redesign both the rubric and the review screen's look. The rubric wins every time, a beautiful screen full of vague items still teaches reviewers nothing.
The model got better, for real: say Garnish's overall accuracy improves next quarter, fewer invented ingredients across the board. That's not a reason to soften the rubric. A rarer mistake still needs a specific, checkable way to catch it when it happens, or it just gets rarer and more surprising.

Where people run it wrong.
They write one item that tries to cover everything, "is this accurate and appealing," which is really two different judgments wearing one checkbox.
They add a five point scale back in "for nuance," and watch agreement drop the exact way it did the first time.
They ask the reviewer to also judge something needing training they don't have, like a legal or medical claim, instead of routing it out.

How to use it live. Say the real tension out loud before answering: "is this asking me for a scoring system, or for a habit I can hand someone with no training." That buys a beat, and it's usually the second one.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
SPARK: design against the failure before you build. Built for design questions, where you're inventing something new rather than reacting to a change.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Aslak Kessing, who runs quality at Garnish Labs, an AI tool that drafts restaurant menu descriptions. He built the review process in the company's first year.
3 · THE HABIT
What habit should the rubric build in the reviewer?
Tap to flip
ANSWER
Answer three or four fixed yes or no facts against the dish's own ticket, instead of rating how the sentence feels.
4 · THE ANCHOR
What's the one design decision the whole rubric hangs on?
Tap to flip
ANSWER
Every item is a checkable fact about this exact output, never a quality judgment. "Does it name an ingredient not on the ticket" is a rubric item. "Rate how appetizing this sounds" is not.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Picking the fuller one to five "how appetizing" scale over a plain approve or reject toggle, because product wanted a number for a future recommender feature, without asking whether it could also catch a real error.
6 · THE NUMBER
Fill in the blank: under the old scale, two reviewers agreed only ___ percent of the time. Under the checklist, agreement rose to ___ percent.
Tap to flip
ANSWER
61 percent, then 94 percent. The catch rate on planted errors moved from 11 of 24 to 22 of 24 in the same swap.
7 · THE REPLAY
Same fake hazelnut line, new rubric, what changes?
Tap to flip
ANSWER
Item one, does this name an ingredient not on the ticket, catches it before publishing, in under thirty seconds, instead of a reviewer giving it a five because it read nicely.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the shared anchor?
Tap to flip
ANSWER
Kindly, an AI tool that drafts shelter animal adoption bios. Same anchor: every rubric item checks a fact against the intake sheet, never a feeling about whether the bio "feels true" to the animal.

Check yourself Score: 0 / 0

True or false
1. True or false: because Garnish's first rubric used a five point scale instead of a single toggle, it was more thorough, so it caught more real errors than a simpler design would have.
  • True
  • False
Show hint
Check the catch rate on the 24 planted errors, not how rich the scale looked on paper.
Show answer
False. It caught only 11 of 24 planted errors, 46 percent, because the scale asked reviewers to rate a feeling rather than check a fact. Richness didn't translate into catching anything.
Multiple choice
2. Why did agreement between reviewers jump from 61 percent to 94 percent after the redesign?
  • A. The reviewers went through a training session on the golden set first.
  • B. Every item became a checkable fact against the dish's own ticket instead of a feeling.
  • C. Garnish Labs hired more reviewers to cross-check each other.
  • D. The model's accuracy improved that same month.
Show hint
Nothing about the model or the reviewers changed. Only the wording of the rubric did.
Show answer
B. Same 150-description golden set, same two reviewers, same model. The only variable that changed was whether the item asked for a fact or a feeling.
Fill in the blank
3. The old five point scale caught ___ of the 24 planted errors in the golden set. The checklist caught ___.
Show hint
It's the same golden set both times, only the rubric changed.
Show answer
11, then 22. Almost double the catch rate, from the exact same descriptions and the exact same two reviewers.
Short answer, name the rejected alternative
4. What alternative did the Garnish Labs team consider when building the first review screen, and why did it lose?
Show hint
Look at the meeting Aslak would take back, eight months before the hazelnut incident.
Show answer
Model answer: A plain approve or reject toggle was on the table early. The team picked the fuller one to five scale instead because product wanted a number that could later train a recommender feature. It lost, eventually, because a number nobody can check is not more informative than a toggle nobody can check, it just felt more sophisticated while catching fewer real errors.
Short answer, apply it yourself
5. Pick a rating you give inside an app you use yourself. Name one place it's really asking for a feeling instead of a checkable fact, and how you'd rewrite it.
Show hint
Think of a star rating or a thumbs up or down that follows an AI-generated result.
Show answer
Model answer: A recipe app asks "how good was this AI-suggested substitution, 1 to 5 stars," after I swap an ingredient. That's a feeling. I'd rewrite it as "did the substitute change the dish's texture in a way the recipe didn't warn you about, yes or no," something I can actually check against what happened in the pan.
Multiple choice
6. The rubric's real production bar isn't zero invented ingredients, ever. What is it instead?
  • A. A new rubric version has to clear at least 90 percent catch rate on the golden set before it ships.
  • B. Garnish must never invent an ingredient again after this fix, full stop.
  • C. Reviewers must give every description a score above 4.
  • D. The model gets retrained from scratch every time one description gets flagged.
Show hint
A model answering from a short ticket will occasionally embellish. The bar isn't about the model never being wrong.
Show answer
A. The bar is probabilistic, not a promise of zero mistakes. A rubric version has to clear 90 percent catch rate on the golden set, the same kind of gate a new model version has to clear before rollout.
Before you close the answer
Why this works
Tests whether you can turn "design a rubric" into one concrete, checkable design decision instead of a wish for good judgment. Most candidates describe a quality process. Few say what makes a single item actually checkable by a stranger with no training.
Follow-up traps
"Isn't a five point scale just more information than yes or no?" Response: more options isn't more information if two honest reviewers can't agree on what a 3 means versus a 4. The scale added range, not signal, and the golden set proved it, 61 percent agreement.

"What if the reviewer disagrees with what counts as 'on the ticket'?" Response: that's exactly why the ticket, the dish's own short ingredient list, is the fixed source of truth for every item, not the reviewer's memory or the finished sentence.
If pressed
Garnish Labs versions its rubric the same way it versions the model. A new rubric item doesn't go live for every restaurant at once, it runs against the golden set first, and only ships once it clears the 90 percent catch rate bar, the same gate a new model version has to clear before rollout.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more