CaseAdvancedQuality, Cost & Token Economics / Success metrics for AI products / #12

How do you measure quality when users rarely give explicit feedback?

What I'd actually build
Watch what the photographer does with the edit, not what they say about it. Track three moves only: kept as is, thrown out, or heavily reworked, and score each one against that one photographer's own normal habits, not one fixed rule for everyone. Treat a rework that breaks from a person's own pattern as the real warning sign, because plenty of photographers rework everything out of habit, and that has nothing to do with the model.
Build this, in order
  1. Watch behavior instead of waiting for a rating.Why: no rating button exists on an editing tool, and forcing one just adds a task nobody wants mid edit.
  2. Track three moves only: kept, discarded, heavily reworked.Why: these are the only signals already sitting in the product, visible with no extra tap from anyone.
  3. Score every edit against that one photographer's own 90-day habit, not one fixed cut off for everyone.Why: a veteran retoucher reworks nearly everything by habit, and a flat rule would flag her every single day for no reason.
  4. Ship a model version only when its baselined kept rate clears a calibrated band on a hand-graded set of photos.Why: one good week proves nothing. A real bar needs a set senior retouchers already scored.
  5. Re-check the signal itself against real human grading every few weeks.Why: a number that quietly stops meaning what it used to is worse than no number, because nobody thinks to doubt it while it is still moving.
  6. Leave reason codes and per-part attribution for later.Why: guessing why someone discarded an edit, before the kept versus discarded signal itself is proven, just adds noise on top of an unproven number.

How to answer this, stage by stage

Nobody is grading whether you can name a signal. They are grading whether you will trust a number you cannot verify, or build one you can. Six moves get you there.

1
Ground it in one real product, one real photographer
Say it like this
"Let's make this real. Loomstone makes Petalgrain, an app that takes a photographer's photo and runs one auto pass: it fixes the exposure, smooths out blemishes, and cleans up the background. Dova Onyekachi is the quality lead there, and every Friday she has to tell the team whether the newest retouch model is actually better."
Why this works
Grounds the whole answer in a real product and a real job before naming a single signal.
2
Say the real question out loud
Say it like this
"Here's what this question is actually asking. It's not how do I get more ratings. It's how do I read quality off behavior that already exists, since a rating button is never going to get built into an editing flow people want to move fast through."
Why this works
Reframes the question from a wish, get more feedback, into a design problem, read the feedback that already exists.
3
Name the habit you want the system to build in her
Say it like this
"The habit I want Dova to build is simple. Stop waiting on a number that will never arrive. Start reading what the photographer actually does next: keep it, bin it, or fight it."
Why this works
This is the payoff step. It names what changes in the person, not just what gets built for them.
4
Give the anchor, the one decision everything else hangs on
Say it like this
"Here's the anchor. Track kept, discarded, and heavily reworked, and score every one of those against that photographer's own last 90 days, not one fixed number for the whole app. A wedding photographer who reworks every single photo out of habit is not the same signal as a hobbyist who suddenly starts reworking every photo this week."
Why this works
This is the concrete design decision, and it matches the direct answer. A vague "we'd watch usage" fails here.
5
Say what breaks the day that anchor lies to you, and how it survives
Say it like this
"The day this goes wrong looks like a flat topline number. Casual photographers' kept rate quietly fell from 84 percent to 61 percent over a month, and nobody noticed, because it blended into everyone else's number and barely moved. So the fix has to be built to catch that from day one: cut the number by who the photographer actually is, never just the average."
Why this works
Shows you designed against your own stated risk, not just described a risk in the abstract.
6
Close on what you leave alone, and the trade you accept
Say it like this
"I would not try to guess why someone discarded a photo on day one, that's a second system, built later. And I'd accept that checking a photographer's own 90-day habit costs more to store and run than one fixed rule for everyone. That's the trade. A cheaper number that lies to Dova about her best users every single day is not actually cheaper."
Why this works
Naming a deliberate cut and a real cost is what makes this a decision, not a wish list.
If you remember one thing A number that barely moves is not proof nothing is wrong. It is proof you have not yet asked whether it is the same number for everyone inside it.

Let's learn

Petalgrain is a button. A photographer drops in a photo, taps auto pass, and Petalgrain fixes the exposure, retouches the skin, and cleans up the background, all in a few seconds. No sliders, unless the photographer wants them.

Hand sketched scene titled How Dova checks quality today, without a signal system. Left panel, a person scrolling through 40 exported photos by hand, guessing which ones look off. Right panel, a document labelled the quarterly survey, 4 percent reply rate, answers arrive weeks late.
Before any behavior signal existed, this was Dova's whole read on quality: a manual scroll and a survey that mostly stayed empty.

Before Dova had any real signal, quality reporting meant that quarterly survey. It reached about 4 percent of active photographers, and the answers came back six to eight weeks after a model had already shipped, updated, and shipped again.

Petalgrain's engineers built something better: a kept, discarded, or reworked count, live within two days of a new model going out. In its first month, the newest retouch model, version 14, held a kept rate around 71 percent, barely different from version 13's 69. Leadership called it a quiet win and moved on to the next thing.

Hand sketched labelled parts diagram titled The anchor, what the photographer does next. Center gauge labelled next move after the AI edit. Four labelled spokes: exports as is, kept. Reverts to original, discarded. Drags sliders hard, reworked. Measured against their own 90 day norm.
The anchor is this whole picture at once: three moves, always read against the one person's own habit, never a flat rule for the whole app.
Four weeks, two numbers, only one of them told the truth
Blended kept rate, all photographers
72% 70% 68% week 0 week 4
Kept rate, casual photographers only
84% 72% 61% week 0 week 4
The blended line wobbled inside its normal range the whole time. It was never going to be the number that caught this.

Here is the turn. That 71 percent was true, and it was also hiding something. Casual photographers, who make up most of Petalgrain's users, saw their own kept rate slide from 84 percent down to 61 percent over the next month, a real drop, while the blended number for everyone barely moved, sitting between 68 and 72 the whole time.

We didn't lose 4 points. Casual photographers lost 23, and the average never said so.

Version 15 had started blowing out highlights on bright outdoor shots, exactly the kind casual photographers shoot most. Veteran retouchers barely noticed, because they already rework almost every photo out of habit, exposure blowout or not. Their number stayed close to where it always sat, and it quietly dragged the blended average back toward normal, so nothing on the topline dashboard looked wrong.

Knowledge spark: what is a baselined rate? A rate measured against one person's own normal, not one fixed number for everyone. A 90 percent kept rate might be great for one photographer and a warning sign for another, if that second one is used to keeping 99.
Hand sketched comparison titled Same heavy edit, read against two different baselines. Left panel, a person labelled Reyna new user, caption heavy edit breaks from her calm norm, real signal, AI missed the shadows. Right panel, a person labelled Tomas veteran retoucher, caption heavy edit is his norm, baseline catches that, no false alarm.
Same action, a heavy rework, means two completely different things depending on whose habit it's measured against.

At its worst, this cost more than a bad week of exposure. Petalgrain's biggest wedding photography client, who shoots almost entirely outdoors, quietly stopped using auto pass on outdoor sets for 11 days before anyone at Loomstone knew why. She didn't file a support ticket. She just started editing by hand again, the way she had before Petalgrain existed.

The choice I would take back is the one made at launch: one fixed cut off for what counts as a heavy rework, the same number for every photographer. That was a sensible default when almost every user was a casual weekend shooter doing roughly the same amount of editing. It stopped being sensible the day veteran retouchers, who rework nearly everything by habit, made up a real share of the app's users.

What I would leave alone: the background cleanup part of auto pass. Casual and veteran photographers keep that part almost every time, over 98 percent, no matter who they are. Building a personal baseline for a part of the product that already works the same for everyone would just be more machinery with nothing to catch.

The lesson: a number that barely moves is not proof nothing is wrong. It is proof you have not yet asked whether it is the same number for everyone inside it.

Now here is the same thing as a story

The short version sits above. Read on for how ordinary the week this almost went unnoticed felt from inside Loomstone.

Dova Onyekachi has run quality numbers for six years, three of them at Loomstone. She inherited Petalgrain's whole quality report from a data scientist who left eight months back, and for a while that meant learning the job from old spreadsheets nobody had touched since.

For most of that year, the kept rate did exactly what a good number should do. It climbed a little, quarter over quarter, and every product review Dova sat in ended the same way: numbers up, ship the next version.

Version 15 went out on a Tuesday. By Friday, the blended kept rate sat at 69 percent, one point under the week before. Dova noted it, called it normal wobble, and closed the spreadsheet a little after eleven that night.

The topline number never once looked wrong. It just quietly stopped being true for the people it mattered most for.

Eleven days later, Petalgrain's ops lead pulled ten photos at random from that week's kept pile, the way she did every so often, mostly to keep everyone honest. Two of the ten were outdoor wedding shots with visibly blown out skies, the kind of mistake nobody would call kept on a second look.

That pulled Dova back into the data, this time cut by who the photographer actually was. Casual photographers, the biggest group on Petalgrain by far, had fallen from an 84 percent kept rate to 61 over exactly those four weeks. Veteran retouchers, who rework almost every photo whether it's good or not, had barely moved, and their steady number had been quietly holding the blended average up the whole time.

Two years earlier, when Petalgrain's rework threshold first got built, the design meeting had been short. One fixed cut off, the same for every photographer, felt simple and safe, and almost every user back then was a casual weekend shooter doing roughly the same kind of editing. Nobody in that room was picturing a professional retoucher who reworks 90 percent of everything by habit. Why would they. Loomstone barely had any yet.

What Dova actually did: she rebuilt the kept rate to run against each photographer's own last 90 days, not one number for the whole app, and she pulled a hand-graded set of 300 photos, scored by three senior retouchers, to set the real bar version 15 had to clear before anyone called it fixed. Checked this way, version 15 came in well outside that bar. It got pulled back for casual users within a week, patched, and re-tested against the same 300 photos before it shipped again.

The replay: four weeks after the patch, casual photographers' kept rate was back to 82 percent, close to where it started, and this time Dova had a number that would have caught the drop on day three instead of day eleven.

What I would tell myself, back before any of this: a number that keeps climbing for everyone at once was never going to warn you. You have to go build the one that can climb for one group while it falls for another, and still show you both.

SPARK, five decisions in order

This is a design question, build the measurement system before the flaw shows up, so SPARK fits. Not a diagnosis of something that already broke, and not a ranking of what to build first.

S
Situation. Where the job gets done today, without the fix.
Dova's only read on quality, before this, was a quarterly survey that reached 4 percent of photographers, six to eight weeks late.
In this story: Dova's Friday spreadsheet, before the kept signal existed.
P
Payoff. The habit you want the system to build.
Stop waiting on a rating. Start reading kept, discarded, and reworked as the real answer, since that's already sitting in the product.
Every photographer already answers this question just by using the app.
A
Anchor. The one design decision everything else hangs on.
Track kept, discarded, and heavily reworked, scored against each photographer's own last 90 days, never one fixed number for everyone.
This is the direct answer. Everything else in this system exists to protect it.
R
Risk. What breaks the first time you're wrong.
A photographer who reworks almost every photo by habit looks exactly like a photographer rejecting a broken model. The baseline is what tells them apart.
Casual users' real regression hid inside the blended average for eleven days before an audit caught it.
K
Keep out. What you deliberately do not build on day one.
No reason codes, no guess at why someone discarded a photo, no per-part attribution splitting exposure from retouch from background cleanup. Just the three-way signal, baselined, first.
Guessing why on top of an unproven number just adds noise to noise.

Two things worth naming straight, since this is where the real judgment sits. Loomstone looked at bolting on a one-tap rating after every edit, a simple thumbs up or down. It got turned down on purpose: forcing a rating mid edit breaks the flow the whole app is built around, and the photographers who'd actually tap it are the ones who already feel strongly, which skews the number toward complaints and skips everyone quietly satisfied. The failure worth naming by name is a signal that quietly stops meaning what it used to, the kept rate could stay perfectly healthy even while a model gets subtly worse, if the flaw is one a photographer just doesn't notice on a phone screen. The guardrail is a standing check, not a one time build: every week, a senior retoucher hand grades 40 of that week's kept photos against the golden set, checking that kept still means good, not just unnoticed. And the bar itself isn't a single clean number. A retouch model version clears the gate when its baselined kept rate stays inside a calibrated band of the version before it, across a rolling two-week window, checked against those 300 hand-graded photos, not one good day. The trade is real: baselining every photographer's own habit costs more to store and compute than one fixed rule, and waiting ten minutes after export to call an edit truly kept, instead of counting it the second it happens, keeps the dashboard a little behind the actual afternoon. That delay is the price of a number that tells the truth instead of one that's just fast.

And if you want to be sure it really works, try it somewhere else

Same five letters, a veterinary tool instead of a photo app, so the method proves itself instead of repeating a story I happened to prepare.

Briarcroft runs Notewell, an app that drafts a vet's visit notes from the audio of an exam. Devan Brenner runs product for it.

S, situation. Vets today write every note by hand after each appointment, or dictate into a recorder and type it up later that night, whichever the clinic happens to have set up.
P, payoff. Stop asking vets to rate a note. Watch whether they keep it as dictated, lightly edit it, or discard it and start over.
A, anchor. Track kept, edited, and discarded, baselined against each vet's own last 90 days of notes, not one fixed edit rate for the whole clinic.
R, risk. One emergency vet on the night shift dictates very short, clinical notes, then always adds a warm, plain-language line for the pet owner before saving, every single time, good AI draft or not. Read flat, her near-constant edit looked like a constant complaint about Notewell's writing. Baselined against her own habit, it wasn't a signal at all.
K, keep out. No attempt yet to tell a real clinical correction apart from a stylistic add. That split comes later, once the keep versus edit signal itself is trusted.

False quality alarms on one vet's notes, before and after baselining
89% 6% Flat edit rate rule Baselined to her own norm 89% 6%
Almost every one of her notes tripped the flat rule. Measured against her own habit, nearly none of them did, and the real complaint days still stood out.
Same shape, different stakes At Petalgrain, the unwatched cost was a wedding client quietly going back to editing by hand. At Briarcroft, it's a real complaint about a medicine mix-up getting buried inside a vet's normal habit of rewriting every note anyway.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the anchor: whatever the product is, track kept, discarded, and reworked, baselined per person, never one number for everyone.
Cost: engineering says the personal baseline can't ship for two months. Don't fall back to one fixed cut off in the meantime and call it good enough. Hold the riskiest model versions back from wide release until the baseline exists.
The model got better, for real: say the retouch model gets meaningfully sharper at exposure. That's still not proof a flat cut off is safe. A better model just makes the next hidden regression harder to spot without the baseline.

Where people run it wrong.
They treat a flat topline number as proof nothing broke, instead of asking whose number is flat and whose actually moved.
They read a high edit rate as rejection across the board, instead of checking whether that person edits everything on principle.
They build the personal baseline once and never check it against real human grading again, so a signal that quietly stopped meaning anything keeps getting trusted for months.

How to use it live. Say the split before naming a single fix: "I always ask whether the behavior I'm reading means the same thing for every person doing it, because the same action, a rework, an edit, can be a complaint from one person and a habit from another." That buys you room to give the real answer instead of reaching for "add a feedback button" on reflex.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What framework fits designing a quality signal when there's no explicit feedback, and why?
Tap to flip
ANSWER
SPARK. You're designing the measurement system itself before it fully exists, not diagnosing something that already broke.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Dova Onyekachi, Loomstone's quality lead for Petalgrain, six years into analytics work and the only person who owns this report.
3 · THE HABIT
What habit was the new system built to create in her?
Tap to flip
ANSWER
Stop waiting for a rating that will never come. Read kept, discarded, and reworked as the real answer, since it's already sitting in the product.
4 · THE ANCHOR
What's the anchor, the one decision this whole answer hangs on?
Tap to flip
ANSWER
Track kept, discarded, and heavily reworked, scored against each photographer's own last 90 days, never one fixed number for everyone.
5 · THE OLD DECISION
What old decision does this answer take back?
Tap to flip
ANSWER
The launch-era choice of one fixed rework cut off for every photographer. It made sense when almost everyone was a casual shooter; it stopped making sense once veteran retouchers made up a real share of users.
6 · THE NUMBER
Fill in the blank: casual photographers' kept rate fell from 84 percent to ___ percent over four weeks, while the blended average barely moved.
Tap to flip
ANSWER
61 percent. The blended number wobbled between 68 and 72 the whole time and never gave it away.
7 · THE REPLAY
Same audit day, new design, what changes?
Tap to flip
ANSWER
The baselined signal catches the casual regression by day three instead of day eleven, and version 15 gets pulled and patched before it reaches most outdoor shooters.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's its anchor?
Tap to flip
ANSWER
Notewell, Briarcroft's veterinary note-drafting tool. Its anchor is the same kept, edited, discarded signal, baselined per vet, so one vet's habit of always adding a warm client note doesn't get misread as a complaint.

Check yourself Score: 0 / 0

Multiple choice
1. What is the anchor decision in this answer?
  • A. Add a one-tap rating after every edit.
  • B. Track kept, discarded, and heavily reworked, scored against each photographer's own last 90 days.
  • C. Ask photographers to fill out a short survey once a month.
  • D. Count how many photos each photographer exports per week.
Show hint
It has to work without ever asking anyone anything.
Show answer
B. This is the only signal already sitting in the product, and it survives the personal-style problem because it's measured against each person's own habit.
Fill in the blank
2. Casual photographers' kept rate fell from 84 percent to ___ percent over four weeks, while the blended average barely moved.
Show hint
Check the chart in "Let's learn."
Show answer
61 percent. The blended number moved four points the whole time. The real number, for the group that mattered most, moved 23.
True or false
3. True or false: a blended kept rate that barely moves proves the newest retouch model is safe for every kind of photographer.
  • True
  • False
Show hint
Check who Petalgrain's biggest user group actually is.
Show answer
False. The blended number can sit still while one whole group's real number is falling, exactly what happened to casual photographers on version 15.
Short answer
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look for the design meeting memory, not a setting anyone could just turn up.
Show answer
Model answer: The one fixed rework cut off for every photographer, set when Petalgrain first launched. It made sense because nearly every user back then was a casual weekend shooter with a similar editing pattern. It stopped making sense once veteran retouchers, who rework almost everything by habit, made up a real share of the app's users.
Short answer, apply it yourself
5. Think of an app you use that has no rating button anywhere in it. What behavior would tell that app whether you were actually happy, without ever asking you?
Show hint
Look for the thing you do right after the app hands you something, not a rating.
Show answer
Model answer: A grocery delivery app that swaps out unavailable items for me. Whether I keep the substitute in my cart or delete it the moment I see it is a cleaner signal than any star rating, and it's already sitting in the data with no extra tap from me.
Short answer, work it out
6. If casual photographers had made up an even bigger share of Petalgrain's users that month, would the blended average have hidden the regression better or worse? Why?
Show hint
Think about what a blended average actually does to a bigger group's number.
Show answer
Model answer: Worse at hiding it, not better. A blended average leans toward whichever group makes up more of it. A bigger casual share would have pulled the blended number down further, closer to 61 percent, so the topline number would have looked far more obviously wrong, not less.
Before you close the answer
Why this works
Tests whether you'll trust a topline number just because it's calm, or go find out whether it's calm for everyone underneath it. Most candidates stop at "add a feedback button."
Follow-up traps
"Isn't scoring every photographer against their own baseline just going to hide a model that's bad for absolutely everyone?" Response: no, because a model that's bad for everyone drags every single person's own recent baseline down at the same time, and that still shows up as each person falling below their own normal, no cross-user comparison needed to catch it.

"Doesn't waiting ten minutes after export before counting something as kept just let a bad model slip through in the meantime?" Response: the wait only delays counting, not detection, since it's applied evenly to every edit. A bad model still gets caught inside the same rolling two-week window, it just means each individual "kept" label is trustworthy instead of an early guess a photographer changes their mind about five minutes later.
If pressed
The golden set itself gets refreshed. Three senior retouchers rescore 60 new photos into it every quarter and retire the oldest 60, so the bar a new model has to clear never quietly goes stale against photos years out of date with how people actually shoot today.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more