Describe how you would measure quality for a creative generation feature.
Quality on a creative feature is not a score you compute. It is a question you ask real people, blind, against a real alternative, before you trust anything the model says about itself.
- Build a blind head-to-head panel: real small business owners compare the AI logo against a human-made one on the same brief.Why: it's the only test that touches what "good" actually means here, whether a stranger would pick it.
- Stop treating the checklist rubric as the finish line.Why: skip this and a logo that passes every box but looks like clip art ships anyway.
- Set the ship bar as a win rate, not a percent.Why: "clears 55 percent of blind picks against a human alternative, on a rolling 30-test panel" is a real bar. A single pass or fail number isn't.
- Keep the rubric, but demote it to a cheap first filter.Why: it still catches broken output fast and free, illegible text, clashing colors, before anything reaches the slower, costlier panel.
- Don't hand the grading job to a second AI model as the main signal.Why: a model trained on similar work shares the same blind spots as the one it's grading, so it keeps calling safe, forgettable output good.
- Skip the panel on parts that are just resizes of an already-approved logo.Why: a favicon pulled from a logo that already won its panel isn't a new creative decision, the cheap check is enough there.
How to answer this, stage by stage
Nobody's grading whether you know quality matters. They're grading whether you'd trust a machine to grade its own homework, or put a real person in the room instead.
Let's learn
What does a checklist actually know about whether a logo is any good?
Emblynn is a tool that builds a small business a logo and a starter brand kit, a color palette, two font pairings, one page showing how to use them, from just a business name and a short description of what the business does.
Before a tool like this existed, a small business owner hiring a freelance designer for the same kind of kit paid around $400 and waited two to three weeks for a first draft. Emblynn builds one in about 90 seconds.
To keep quality from slipping as that pipeline scaled, the team built a rubric: twelve checklist items, contrast ratio, font legibility, no clipped text, colors that match the brief's mood words, each logo auto-graded the moment it's generated. Every one of the roughly 1,400 kits made each week gets scored in under a minute, no person involved. For months the average sat at 96 percent. Steady. Green across the dashboard.
Here's the part that matters. A 96 percent average is not proof the logos were good. It's proof they mostly weren't broken. Those are not the same claim, and nobody had checked whether they lined up.
What that costs at its worst: a bakery owner pays Emblynn's $29 a month, gets a kit that scores 98 percent, puts it on her shopfront sign and her takeout bags. Three weeks later a regular tells her it looks like every third bakery on her Instagram feed. She pays a freelance designer $400 to redo it anyway. She's out the subscription, the printing costs, and the three weeks she could have spent finding a designer directly. Emblynn didn't just fail to help her. It cost her more than doing nothing would have.
What I would leave alone: the mechanical resizes generated from a logo that's already won its panel test, the favicon, the letterhead, the social media avatar, don't need their own comparison. They're not new creative decisions, they're the same approved mark at a different size. The cheap rubric check, does it stay legible that small, is genuinely enough there. Running every resize through a full panel would spend real time and money on cases where nothing creative is actually being judged.
The lesson: a checklist can prove an output isn't broken. It can't prove a stranger would choose it over something a person made. For anything creative, that second question was always the real one, and it was never going to answer itself.
Now here is the same thing as a story
Read this when you want to feel why the gap mattered, not just know that it existed.
Emblynn's office runs quiet on Monday mornings, except for one screen everyone glances at on the way to the coffee machine: the weekend's quality dashboard.
Iolanthe Belasco has run quality there for four years. For most of that time she could read that dashboard in ten seconds and tell you if the week was healthy, the way a mechanic hears an engine and already knows. When the rubric first shipped, she still pulled twenty logos a week and looked at them herself, just to see if the checklist and her own eye agreed. They always did. So she cut it to ten. Then five. By the second year she wasn't opening any of them. The dashboard said 96 percent. What was there to look at?
Then, on an ordinary Monday, Silke Novgren, a senior designer who'd joined the team six months earlier, pulled up that week's top-scoring kit, a 98, on one half of her screen. On the other half she pulled up what a freelance designer, working off the exact same brief, had actually delivered to a different client the year before. She turned her monitor toward Iolanthe.
"This would pass," Silke said, pointing at the 98. "So would a parking sign."
Iolanthe looked at both for a long moment. The AI kit wasn't wrong about anything. The colors matched the brief. The type was legible. It also looked like six other cafe logos she'd scrolled past that same morning without remembering a single one of them. The human version had a small, specific choice in it, an odd little mark worked into the letterform, that she kept looking back at.
She didn't argue with Silke. She asked for three days to find out if this was one bad example or the whole picture.
The three days weren't spent proving the rubric wrong about anything it actually checks. They were spent building the test nobody had ever run: thirty briefs, each one given to Emblynn and, blind, to a real freelance designer, then shown side by side, unlabeled, to actual small business owners in similar trades, and asking one plain question. Which one would you use for your own shop?
The AI logo won ten times out of thirty. Thirty-four percent. Against a 96 percent rubric average on the exact same kits.
The old decision that opened this gap went back two years, to a meeting about review backlog. Reviewers were drowning, and someone asked whether every kit really needed a human look before it shipped. The answer, reasonable at the time, was no: keep a person checking anything under 90 percent, and let anything at or above ship straight through, since the rubric had never let anything genuinely broken past that line. Nobody in that room asked whether 90 percent correlated with anyone actually wanting the logo. It correlated with not being broken. That was the number they had, so that was the number they used.
Run the same Monday again with the panel already built. The kit that would have scored 98 and shipped in ninety seconds instead goes into that week's rolling comparison first. It loses, two picks out of five. The team doesn't retrain Emblynn from scratch, they use what the panel actually flagged, safe layouts, safe type pairings, the same three color moods over and over, to change what the generator gets rewarded for showing. By week eight, kits from that same kind of brief are winning three picks out of five. Fifty-eight percent, past the line where a design ships without a second look.
One design trusted a machine to grade its own homework. The other put an actual customer in the room before anything shipped.
What I'd tell myself, back in that backlog meeting: a threshold chosen because it was cheap to compute is not the same thing as a threshold chosen because it's the thing that actually matters. We had the first kind for two years and called it the second.
SPARK, the design behind the number
This isn't a bug fix on a broken feature. It's SPARK run on a question that only sounds like a metric question, using a panel of real people as the thing being designed, not the thing being measured.
Three things worth stating directly, since this is where the real judgment sits. The rejected alternative was padding the existing rubric with a thirteenth checklist item for "originality" or "distinctiveness." It lost because a checklist item for being surprising just becomes one more box the model learns to tick in the safest, most predictable way possible, which is the exact failure the whole redesign exists to fix. The AI-specific failure mode worth naming is a kind of evaluator collusion: a second AI model trained on similar creative work would share the first model's blind spots and keep rating safe, on-distribution output as good, since neither model has any actual stake in whether a stranger finds the result memorable. The guardrail is the panel itself, real people with no relationship to Emblynn, rotated each round so the same five tastes don't quietly become the new checklist. That guardrail isn't free: a blind panel costs real money and time against the rubric's free, instant check, roughly $45 a comparison, a $35 freelancer fee plus a small panel incentive, and a three day turnaround against the rubric's under-a-minute pass. Emblynn accepts that cost only where it's load-bearing, thirty comparisons on a rolling two-week sample, not on every one of the 1,400 kits made that week. And the bar for shipping without another look isn't a perfect score, a creative model can't promise that and shouldn't be asked to. It's a win rate that clears 55 percent against a human-made alternative on that rolling panel, checked every two weeks, not a single pass or fail number pretending to be certain about something inherently a matter of taste.
And if you want to be sure it really works, try it somewhere else
Same five letters, a jingle generator for local radio ads instead of a logo tool, with nothing visual anywhere in sight.
Chorusmint writes and sings a fifteen second jingle for a local radio ad from a business's name and a short brief. Wynter Renfield runs quality there.
S, situation. Today, without a panel, quality is an auto-score: does the jingle fit the ad slot's exact length, is the business name spoken clearly, is the audio clean. Every jingle passes or fails on those three checks alone.
P, payoff. The habit to build: compare the AI jingle against what a real jingle writer made for the same brief, blind, instead of trusting a technical pass.
A, anchor. Play both jingles, unlabeled, to actual local shop owners in the same trade, and ask which one they'd want playing under their own business name on the radio.
R, risk. A jingle that hits every technical box, clean audio, correct tempo, a clearly spoken name, but has a melody nobody hums after hearing it once. A checklist has no way to check for memorable, and memorable is the entire point of a jingle.
K, keep out. No AI model "listening" to the first model's jingle and scoring it for catchiness as the primary signal. Two models that have never heard a real ad break will agree with each other for reasons that have nothing to do with what makes a jingle stick.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the anchor and the one gap number, don't spend the time narrating the checklist that didn't work.
Cost: there's no budget yet for a formal panel. Don't skip the check, run it with five people from outside the design team instead of a paid panel until the budget exists.
The model got better, for real: say Emblynn's rubric score climbs to 99 percent this quarter. That's not proof the creative gap closed. A model that improves on average can still keep losing to a human on the one thing a checklist was never built to see.
Where people run it wrong.
They read a high rubric score as proof there's nothing to fix, and never build a panel at all.
They build the panel once, get a good number, and never run it again, so a later change to the generator ships completely unchecked.
They let an AI model grade the comparison for speed, and it quietly prefers its own kind of output every time.
How to use it live. Say the two questions are different before answering either: "does it pass a checklist, or would someone actually pick it." That line buys you a beat to think instead of reaching for "accuracy" out of habit.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Why not just have the model grade its own output and skip the human panel entirely?" Response: because a model grading similar work shares the same blind spots as the model that made it. Two systems trained on the same kind of data agreeing with each other proves nothing about whether an actual customer would pick it.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Quality metrics: accuracy vs usefulness vs trust
- #1 Define accuracy, usefulness and trust as three distinct measurable properties.
- #2 Give an example of an output that is accurate but not useful.
- #3 Give an example of a product that is useful despite being frequently wrong.
- #4 How would you measure trust in an AI feature?
- #5 Explain why improving accuracy can decrease trust.
- #6 Describe the calibration problem: what happens when confidence does not match correctness?