ConceptIntermediateQuality, Cost & Token Economics / Quality metrics: accuracy vs usefulness vs trust / #3

Give an example of a product that is useful despite being frequently wrong.

Pick a side first, no hedging. Then find the one kind of wrong that would actually sink the product, and say what number would prove it's happening.

The direct answer
Pick a product where being wrong is cheap and fast to catch, and where the win when it's right is worth more than the cost of catching it when it's not. Glintbox, an AI tool that throws marketers dozens of rough campaign concepts and taglines to react to, gets more than half of every batch wrong and is still worth keeping open, because a bad concept costs three seconds to skip. The design only stays defensible as long as the one kind of wrong that isn't cheap, a made up number or claim dressed up as a fact, gets caught before a person ever acts on it.
Do this, in order
  1. Keep Glintbox fast and wrong often. Don't slow the whole thing down to be right more.Why: a bad concept costs three seconds to skip, and a real spark can carry a whole campaign. Slowing everything down to catch the cheap mistakes costs more than it saves.
  2. Guard only the one kind of wrong that's actually expensive: a made up claim dressed up as a fact.Why: that's the kind that's hidden and costly. Tone misses and flat angles are free, everyone already shrugs those off.
  3. Name who pays for each kind of wrong, in real seconds or real dollars, not percentages.Why: a percentage sounds like a grade. A number tied to a real person's real minute is what makes the pick defensible.
  4. Set a real kill number now, before anyone asks for one.Why: a pick with no number that would flip it is just a preference wearing a decision's clothes.
  5. Leave every other kind of wrong alone.Why: fixing mistakes nobody minds wastes the budget that should go toward the one kind that actually costs something.
  6. Watch for the moment a person stops checking at all, not just for the error rate going up.Why: the whole design leans on someone standing between the tool and the world. If that person quietly stops showing up, the design's real assumption just broke.

How to answer this, stage by stage

Nobody is grading whether you can name a product. They're grading whether you can defend picking one side of a tradeoff and then tell them exactly what would make you change your mind. Six moves get you there.

1
Ground it in one real product before you say anything abstract
Say it like this
"Let me make this concrete. Glintbox is a tool a marketer types a brief into, and it hands back forty rough campaign concepts and taglines in about twelve seconds. Seraphina, an associate creative director, opens it most Monday mornings."
Why this works
A tradeoff argued in the abstract is just an opinion. Tied to one real product and one real person, it becomes something you can defend under follow up.
2
State your position first, with no hedge
Say it like this
"My answer is yes, and I'll go further: I'd rather it stay wrong most of the time than get slower and safer. It's more useful wrong and fast than right and careful, for the specific job it's doing."
Why this works
Interviewers are checking whether you'll actually commit. "It depends" is the answer that fails this question every time.
3
Name who feels each kind of wrong, in real numbers
Say it like this
"Say the model is wrong on twenty two of the forty concepts. Seraphina skims each one in about three seconds and moves on. That's a minute of her morning. It's already priced into the four minutes she budgets for the whole batch."
Why this works
Turning a percentage into a real cost, priced in seconds, is what separates a real judgment call from reciting a metric.
4
Find the one asymmetry the whole pick actually rests on
Say it like this
"Not all wrong is equal here. A bad angle is cheap, she never trusted any single one of them anyway. A made up statistic sitting inside a good sounding line is a different animal, because a specific number reads as more true, not less, and nobody's fact checking every line at brainstorm speed."
Why this works
This is the actual answer to a tradeoff question. Both sides always cost something. The job is finding the one that's hidden and expensive.
5
Give the kill criteria before anyone has to ask for one
Say it like this
"I'd flip this pick if the share of lines carrying a real, checkable claim, a number, a name, a quote, climbed past about eight percent of a batch. It sits at two to three percent today. Past eight, a three second skim structurally can't catch them all anymore."
Why this works
A number that would change your mind is what proves you actually reasoned your way here, instead of just landing on a side you liked.
6
Close on the one line that survives a follow up
Say it like this
"So: wrong is fine here, because wrong is cheap here. The design only breaks the day wrong stops being cheap, and I've told you exactly what that day looks like."
Why this works
Ending on the pick, restated, plus its own failure condition, leaves nothing for the interviewer to catch you not having thought about.
The line to fall back on Wrong often is fine. Wrong and unlabeled is not. That second sentence is the part of this answer worth defending hardest.

Let's learn

Every big pitch used to start the same way: Seraphina alone, a blank document, and most of an afternoon gone before she had even three directions worth showing anyone. On a hard brief, closer to three hours.

Glintbox is an AI tool a marketer types a brief into. It hands back a stack of forty rough campaign concepts and taglines in about twelve seconds. Most of them are, honestly, not very good.

Now Seraphina types the brief in, waits about twelve seconds, and skims all forty. Roughly three seconds each, tossing anything off brief, dull, or just wrong for the brand. In a bit over two minutes she usually has three or four worth developing further.

Hand sketched two panel comparison titled Two kinds of wrong, not one. Left panel, small and plain, labeled Wrong tone or angle, caption skimmed, dropped, three seconds gone. Right panel, larger and marked in red orange, labeled A made up fact inside a good line, caption already copied into the deck before anyone checks.
Not all wrong costs the same. One kind gets skimmed and dropped in seconds. The other kind gets believed.

The turn is this: those twenty two extra wrong concepts are not the real problem. Seraphina was always going to throw most of the batch away. The real question is what she does with the ones that sound right but aren't.

Forty guesses that cost three seconds each are not the risk. One guess that reads like a fact is.
Why does the model invent a number at all? It isn't looking anything up. It's guessing the next likely word, over and over, and a specific sounding number or name usually scores better in that guessing game than a vague one. The model isn't lying on purpose. It just doesn't know the difference between a number it made up and one it read somewhere.

At its worst, this costs a real client relationship. A made up statistic sits in a slide in front of someone who happens to know the real number. Or it ships in a live post, and someone has to write, in public, an explanation of where a fake number came from.

The decision that mattered Early on, the team let the model write specific numbers into taglines freely, because concepts with concrete numbers tested better in early taste tests than vague ones. Nobody tagged which lines carried a real claim and which didn't. It made sense before anyone had actually been burned by one.

What I would leave alone: tone misses, format misses, a joke that just doesn't land. Leave those wild. They're genuinely free to skip, and slowing the whole tool down to catch them wastes the speed that's the entire point of it.

We could have gone the other way and built the careful version first: a mode that checks and sources every claim before showing it to anyone. We tried a rough one. It worked, almost nothing in it was wrong. It also took about five minutes, across several rounds, to hand Seraphina the same three or four keepers the fast, wrong often version gets her in about two.

Minutes to land three or four usable concepts, fast mode vs the careful mode we rejected
6 min 0 2 min 5 min Fast mode, wrong often Careful mode, rejected
What Glintbox actually runsThe accurate alternative we tried and cut
The careful mode only offers five heavily checked concepts per call, so Seraphina needs several rounds to see enough ideas to react against. Fewer wrong guesses, but more than double the time, and far less to actually spark off.

The lesson: the three seconds it costs to skip a bad idea and the phone call it costs to walk back a wrong number in front of a client are not the same size of mistake. A product only earns the right to be wrong constantly if it never quietly lets the second kind through wearing the first kind's clothes.

Now here is the same thing as a story

The short version sits above. Read on for the Thursday a tagline with a number nobody checked almost made it into a client deck.

Seraphina can spot a line that's trying too hard before she's finished reading it. Nine years in agency creative will do that. Give her sixty taglines and she'll sort the real ones from the desperate ones in under a minute, mostly on instinct.

Glintbox arrived at her agency the spring before last. For most of a year it was the best part of her Monday. She'd type in a brief around 9am, get forty concepts back before her coffee cooled, and by 9:15 she'd have three or four real directions circled in the margin.

The habit that quietly changed underneath her happened in three small steps nobody would have called a decision. First, she stopped reading every single concept the way she used to, word by word, because the tool's tone was consistently close enough to brand voice that a careful read stopped feeling necessary. Second, she started treating a concept with a specific number in it as slightly more credible than one without, the same instinct that makes a stat sound truer in a pitch than a vague claim, without noticing she'd started doing it. Third, the product team quietly turned off a small orange flag that used to sit next to any line containing a number or a named claim, because a client testing an early mockup said it made the concepts look unfinished.

Then came a Thursday. Nothing dramatic. A junior copywriter, prepping slides for a pitch the next morning, pulled a Glintbox line straight into the deck: "nine out of ten parents already trust a digital first pediatric brand." It read confident. It read specific. It sounded like something from a real study.

It wasn't from anything. The model had never seen a study. It generated a number because a number like that tends to score well, the same reflex that makes any of its forty lines sound sharper than a vague one.

We didn't lose an afternoon rewriting a slide. We lost the fifteen minutes before the pitch where someone finally asked where that number came from.

Seraphina caught it at 8:40 the next morning, an hour before the client walked in, only because she happened to reread the deck cold instead of trusting last night's version. The line got pulled. The deck survived. Nobody outside the room ever saw it.

I keep coming back to the meeting where the flag got cut. It wasn't a careless call. The taste test genuinely showed the little orange mark made concepts look messier, and at the time nobody had a real incident to point to that said the mark was worth the mess. It made sense then. It stopped making sense the day a junior copywriter, doing exactly what the tool trained her to do, trusted a number because it sounded like one.

Run the same Thursday again with the flag back in place. The line still gets generated. But it renders with a small orange mark and a line underneath: "unverified, check before use." Pulling it into a deck now takes one extra click to confirm someone actually checked it. That click costs about fifteen seconds. It happens on roughly one line out of every forty.

Fifteen seconds against a pitch nearly opening on a number nobody can source. One design hands the room a blank check on how believable something sounds. The other hands it a fifteen second pause, in exactly the one place that pause is worth having.

What I'd tell myself, back in that first meeting: the call to cut the flag was ours, not the copywriter's. We decided fast and confident looked better in a demo than a tool that admits when it's guessing. We got away with it for a year. That's not the same as being right about it.

PICK, in four moves

Not a diagnosis of one bad Thursday. This is PICK run on the actual question: is Glintbox worth keeping even though it's wrong more than half the time, with that Thursday used to test the answer against something real.

P
Position. Say the pick before any of the reasoning.
Yes, Glintbox is worth keeping wrong most of the time. Not "it depends on the client" or "it depends on the brief." A flat yes, with the reasoning to follow, not the other way round.
This is what the direct answer at the top of the page already said in one breath.
I
Impact. Who feels each kind of error, in real units.
The cheap error lands on Seraphina, priced in seconds, already inside the two minutes she budgets for a batch. The expensive error lands on whoever reads the finished line with no idea it came from a guess, a client in a pitch room, or an audience reading a live post, and it costs the agency a scramble, not a shrug.
Naming both sides, not just the one that's easy to defend, is what makes this a real answer instead of a sales pitch for the tool.
C
Cost asymmetry. The one that actually decides the pick.
A wrong tone or a flat angle is visible the second Seraphina reads it, and it costs nothing beyond the three seconds to skip it. A wrong fact hides inside a line that otherwise reads well, because a specific number sounds more true, not less, and nobody independently checks every line at brainstorm speed. Optimize against that one. Leave the rest alone.
Most candidates stop at "there's a tradeoff." The strong answer says which side of it actually bites.
K
Kill criteria. What would flip this pick.
Three lines, none of them a guess. One: the share of concepts carrying a real, checkable claim climbing past about eight percent of a batch, up from a normal two to three percent, since a three second skim structurally can't catch more than a handful. Two: any team shipping Glintbox lines straight out with no editor pass at all, since the entire cost asymmetry assumes a person is standing between the tool and the world. Three: any use inside a regulated category, health claims, finance, anything aimed at kids, where a wrong specific claim is a compliance problem, not an awkward rewrite.
Most sessions should sit near two to three percent. It's not one bad batch that flips the pick, it's a run of batches that don't come back down.
Share of a batch carrying a checkable claim, last six months, against the point that flips the pick
10% 0% kill line, 8% M1 M2 M3 M4 M5 M6
Actual claim rate, monthlyThe line that would flip the pick
Six months, holding between 1.9 and 2.6 percent, nowhere near the eight percent line. That gap is exactly what makes the fast, wrong often design still the right call today, and exactly what a real interviewer is asking you to prove you'd notice if it ever closed.

Worth naming plainly, since this is where the actual judgment sits. The alternative we rejected was routing every concept through the slower, source checked mode by default, the one charted above at five minutes instead of two. It lost because a brainstorming tool's entire value is volume and speed, and a mode built to be right kills both. The AI specific failure worth naming by name is hallucination, the model generating a specific, confident sounding claim it never actually verified, because confident and specific score well in the pattern it learned to produce. The guardrail is a cheap one: a fast pass over every generated line checking for numerals, percent signs, and quoted phrases, flagging anything that matches before the batch ever renders, adding no real time to the twelve second wait. And there's a real cost being accepted here, not a free lunch: catching every claim, always, would mean routing flagged lines through a slower, grounded check before they're ever shown, which is exactly the speed and cost tradeoff the fast mode exists to avoid paying on all forty lines just to protect the one or two that need it.

And if you want to be sure it really works, try it somewhere else

Same four letters, an HVAC dispatch tool instead of a brainstorming one, and the same shape shows up with no marketing brief anywhere in sight.

FaultSketch is an AI tool that reads a customer's phone description of what's wrong with their furnace or air conditioner and guesses the three most likely broken parts before a technician even leaves the shop, so the truck can carry the right part instead of driving back for it. Wenceslas Umeh runs dispatch for a mid sized HVAC company.

P, position. FaultSketch is worth running before every dispatch, even though its own top guess is wrong about which specific part it is more than a third of the time.
I, impact. The cheap error costs almost nothing: the tech still carries the top three guessed parts as a matter of habit, so even a wrong top guess is usually already sitting in the truck. The expensive error is a complete miss, none of the top three guesses is the real fault, and the household waits a full extra day without heat or air conditioning for a second visit.
C, cost asymmetry. A wrong first guess inside the top three costs nothing extra, the part's already on the truck. A complete miss costs a second truck roll, a second missed half day of work for the customer, and a callback that eats the goodwill of the first visit. Optimize against the complete miss, not against the top guess being wrong.
K, kill criteria. FaultSketch's top three guesses cover the real fault about 88 percent of the time today, a complete miss about 12 percent. If the complete miss rate holds above 20 percent for two weeks running, the tool stops paying for itself and needs retraining before it's trusted again. And any call mentioning a safety keyword, gas smell, carbon monoxide, sparking, routes straight to an emergency dispatch regardless of what the model guesses, no threshold, no exception.

Average cost per dispatch, without FaultSketch vs with it
$140 $0 $120 $102 Without FaultSketch With FaultSketch
Base truck roll, $85Callback cost, no tool, 25% miss rateCallback cost, with tool, 12% miss rate
The base cost of sending a truck never changes. What shrinks is the callback tax, from a 25 percent miss rate guessing off the phone call alone, down to 12 percent with FaultSketch's top three guesses on board.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the pick and the kill number: "I'd run it, wrong guess or not, because the cost of wrong is one spare part, not a wasted trip."
Cost: finance asks why keep a tool that's wrong on the specific part over a third of the time. Point at the chart. Wrong on the specific part and still eighteen dollars a dispatch cheaper, because the top three catches almost everything a full miss would cost.
The model got better, for real: say FaultSketch's top one accuracy climbs from 62 to 80 percent next quarter. That's genuinely good news, and it still doesn't retire the kill criteria. A model that's right more often on average can still miss completely on a fault type it rarely sees, and the complete miss rate is the number that would catch that, not the average.

Where people run it wrong.
They watch the top one accuracy number and call it the whole story, when the number that actually protects the business is the complete miss rate.
They let the tool's guess replace the tech's judgment entirely, instead of treating it as a head start the tech still checks against.
They wait for a bad quarter to notice the miss rate crept up, instead of watching it every week the way the kill criteria demands.

How to use it live. Open with the pick and the one number that would flip it, in the same breath: "I'd run it because being wrong costs one spare part, and I'd kill it the day complete misses cross twenty percent for two weeks running." That buys you the room to explain the rest before anyone can push you toward a hedge.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
Which framework fits an "A or B, defend one side" question like this one?
Tap to flip
ANSWER
PICK: Position, your pick in one sentence before any reasoning. Impact, who feels each kind of error. Cost asymmetry, which error is hidden and expensive. Kill criteria, what evidence would flip the pick.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Seraphina, an associate creative director who leans on Glintbox, an AI tool that generates rough campaign concepts and taglines, to spark her Monday morning briefs.
3 · THE HABIT
What did Seraphina quietly stop doing because Glintbox kept being close enough?
Tap to flip
ANSWER
She stopped reading every concept word for word, and started treating a specific sounding number inside a line as slightly more believable than a vague one, without noticing she'd started doing it.
4 · THE TWO KINDS OF WRONG
What are the two kinds of wrong in this story, and which one actually matters?
Tap to flip
ANSWER
A wrong tone or angle, skimmed and dropped in about three seconds, free. A wrong fact, a made up number or claim inside a good sounding line, hidden and expensive. Design against the second one, leave the first one alone.
5 · THE OLD DECISION
What decision does this answer take back, and why did it make sense at the time?
Tap to flip
ANSWER
Cutting the small orange flag that marked any line with a number or named claim, because an early taste test said it made concepts look unfinished. It made sense before anyone had a real incident to weigh against that.
6 · THE NUMBER
Fill in the blank: concepts with a real checkable claim normally sit around ___ percent of a batch. The kill line marked in this answer is ___ percent.
Tap to flip
ANSWER
Two to three percent; eight percent. Below eight, a three second skim can realistically catch what's there. Above it, it structurally can't.
7 · THE REPLAY
Same pitch deck, flag put back, what changes?
Tap to flip
ANSWER
The invented statistic renders with an orange "unverified" mark. Pulling it into a deck now costs about fifteen extra seconds to confirm. That's the price against a pitch nearly opening on a number nobody could source.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs PICK again on a different product. Which one, and what's the parallel?
Tap to flip
ANSWER
FaultSketch, an AI tool that guesses the likely broken part before an HVAC technician is dispatched, run by Wenceslas Umeh. Same shape: wrong about the exact part often, cheap because the truck already carries a backup, expensive only on a complete miss that costs a whole extra day.

Check yourself Score: 0 / 0

Multiple choice
1. Why does it make sense to keep using Glintbox even though more than half of every batch is wrong?
  • A. Because the model will eventually get more accurate on its own.
  • B. Because being wrong here is cheap and fast to catch, and the value of a real spark outweighs that cost.
  • C. Because marketers don't actually read the concepts closely anyway.
  • D. Because forty concepts always contains at least one good one, by definition.
Show hint
Think about what a wrong concept actually costs Seraphina, in seconds, not in principle.
Show answer
B. The whole pick rests on the cost of catching a wrong concept being close to nothing, three seconds, while a real spark can carry a whole campaign. That asymmetry, not the model's overall accuracy, is the actual reason.
True or false
2. True or false: making Glintbox slower and more accurate across every kind of output would be a good fix.
  • True
  • False
Show hint
Look at the chart comparing the fast mode against the careful mode we rejected.
Show answer
False. Slowing everything down to fix the free kind of wrong charges every session for a mistake nobody minded, and more than doubles the time to reach the same three or four keepers, two minutes against five.
Fill in the blank
3. The careful, source checked mode we rejected takes about ___ minutes to reach the same three or four keepers the fast mode reaches in about ___ minutes.
Show hint
Check the bar chart right after the "what I would leave alone" paragraph in "Let's learn."
Show answer
Five minutes; two minutes. The careful mode offers fewer concepts per call, so it takes several rounds to give Seraphina enough to react against, more than double the fast mode's time.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look for the meeting where a small orange mark got cut, in the story section.
Show answer
Model answer: Cutting the orange "unverified" flag on any line with a number or named claim, because an early client mockup test said it made concepts look messier and unfinished. It made sense before there was a real incident to weigh the mess against, and stopped making sense the day a junior copywriter trusted a number because it sounded specific.
Multiple choice
5. Which of these, if it happened constantly, would genuinely NOT need fixing in Glintbox?
  • A. A generated line inventing a specific customer testimonial that never happened.
  • B. A generated tagline that's just flat, clever wordplay for a brand that isn't playful.
  • C. A generated line citing a statistic that isn't from any real source.
  • D. A generated line naming a competitor claim that was never actually made.
Show hint
Ask whether the mistake is a checkable claim or just a matter of taste.
Show answer
B. A tone miss is a taste problem, skimmed and dropped for free. The other three all involve a specific, checkable claim, exactly the kind of wrong this answer says is expensive and worth guarding.
Short answer, apply it yourself
6. Pick an AI product you use yourself that's wrong a lot but still useful. What's the cheap kind of wrong it produces, what would the expensive kind look like, and what number would tell you the expensive kind is creeping up?
Show hint
Ask what you actually do with a wrong output today, and picture the one time that habit would cost you something real.
Show answer
Model answer: A code autocomplete tool. Cheap wrong: it suggests a line that doesn't compile, you delete it, five seconds gone. Expensive wrong: it confidently suggests a line that compiles fine but is subtly insecure, like skipping input validation, and it slips past a quick review because it looks like normal code. The number to watch: how often a suggested line touches security sensitive code, auth, payments, user input, since that's where a quiet miss actually costs something.
Before you close the answer
Why this works
Tests whether you'll actually commit to a side of a tradeoff and then find the one real asymmetry underneath it, instead of listing pros and cons until the interviewer stops you.
Follow-up traps
"What if the client starts treating every Glintbox line as a fact, not just a spark?" Response: that's exactly the kill signal. The moment trust migrates past the person who's supposed to be the check, the flag has to become a hard block, not a suggestion someone can click past.

"Isn't eight percent an arbitrary number?" Response: no, it's picked at the point where a three second skim reliably catches one flagged line a batch and starts missing more than one above it. It's tied to what a person can actually do in the time they actually spend, not chosen to sound tidy.
If pressed
The actual guardrail runs as a fast pass over every generated line before it renders: a check for numerals, percent signs, and quoted phrases. Anything that matches gets the orange flag automatically. It adds no meaningful time to the twelve second wait, because the check is cheap even when the generation isn't.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more