How do you measure quality when users rarely give explicit feedback?
- Watch behavior instead of waiting for a rating.Why: no rating button exists on an editing tool, and forcing one just adds a task nobody wants mid edit.
- Track three moves only: kept, discarded, heavily reworked.Why: these are the only signals already sitting in the product, visible with no extra tap from anyone.
- Score every edit against that one photographer's own 90-day habit, not one fixed cut off for everyone.Why: a veteran retoucher reworks nearly everything by habit, and a flat rule would flag her every single day for no reason.
- Ship a model version only when its baselined kept rate clears a calibrated band on a hand-graded set of photos.Why: one good week proves nothing. A real bar needs a set senior retouchers already scored.
- Re-check the signal itself against real human grading every few weeks.Why: a number that quietly stops meaning what it used to is worse than no number, because nobody thinks to doubt it while it is still moving.
- Leave reason codes and per-part attribution for later.Why: guessing why someone discarded an edit, before the kept versus discarded signal itself is proven, just adds noise on top of an unproven number.
How to answer this, stage by stage
Nobody is grading whether you can name a signal. They are grading whether you will trust a number you cannot verify, or build one you can. Six moves get you there.
Let's learn
Petalgrain is a button. A photographer drops in a photo, taps auto pass, and Petalgrain fixes the exposure, retouches the skin, and cleans up the background, all in a few seconds. No sliders, unless the photographer wants them.
Before Dova had any real signal, quality reporting meant that quarterly survey. It reached about 4 percent of active photographers, and the answers came back six to eight weeks after a model had already shipped, updated, and shipped again.
Petalgrain's engineers built something better: a kept, discarded, or reworked count, live within two days of a new model going out. In its first month, the newest retouch model, version 14, held a kept rate around 71 percent, barely different from version 13's 69. Leadership called it a quiet win and moved on to the next thing.
Here is the turn. That 71 percent was true, and it was also hiding something. Casual photographers, who make up most of Petalgrain's users, saw their own kept rate slide from 84 percent down to 61 percent over the next month, a real drop, while the blended number for everyone barely moved, sitting between 68 and 72 the whole time.
Version 15 had started blowing out highlights on bright outdoor shots, exactly the kind casual photographers shoot most. Veteran retouchers barely noticed, because they already rework almost every photo out of habit, exposure blowout or not. Their number stayed close to where it always sat, and it quietly dragged the blended average back toward normal, so nothing on the topline dashboard looked wrong.
At its worst, this cost more than a bad week of exposure. Petalgrain's biggest wedding photography client, who shoots almost entirely outdoors, quietly stopped using auto pass on outdoor sets for 11 days before anyone at Loomstone knew why. She didn't file a support ticket. She just started editing by hand again, the way she had before Petalgrain existed.
The choice I would take back is the one made at launch: one fixed cut off for what counts as a heavy rework, the same number for every photographer. That was a sensible default when almost every user was a casual weekend shooter doing roughly the same amount of editing. It stopped being sensible the day veteran retouchers, who rework nearly everything by habit, made up a real share of the app's users.
What I would leave alone: the background cleanup part of auto pass. Casual and veteran photographers keep that part almost every time, over 98 percent, no matter who they are. Building a personal baseline for a part of the product that already works the same for everyone would just be more machinery with nothing to catch.
The lesson: a number that barely moves is not proof nothing is wrong. It is proof you have not yet asked whether it is the same number for everyone inside it.
Now here is the same thing as a story
The short version sits above. Read on for how ordinary the week this almost went unnoticed felt from inside Loomstone.
Dova Onyekachi has run quality numbers for six years, three of them at Loomstone. She inherited Petalgrain's whole quality report from a data scientist who left eight months back, and for a while that meant learning the job from old spreadsheets nobody had touched since.
For most of that year, the kept rate did exactly what a good number should do. It climbed a little, quarter over quarter, and every product review Dova sat in ended the same way: numbers up, ship the next version.
Version 15 went out on a Tuesday. By Friday, the blended kept rate sat at 69 percent, one point under the week before. Dova noted it, called it normal wobble, and closed the spreadsheet a little after eleven that night.
Eleven days later, Petalgrain's ops lead pulled ten photos at random from that week's kept pile, the way she did every so often, mostly to keep everyone honest. Two of the ten were outdoor wedding shots with visibly blown out skies, the kind of mistake nobody would call kept on a second look.
That pulled Dova back into the data, this time cut by who the photographer actually was. Casual photographers, the biggest group on Petalgrain by far, had fallen from an 84 percent kept rate to 61 over exactly those four weeks. Veteran retouchers, who rework almost every photo whether it's good or not, had barely moved, and their steady number had been quietly holding the blended average up the whole time.
Two years earlier, when Petalgrain's rework threshold first got built, the design meeting had been short. One fixed cut off, the same for every photographer, felt simple and safe, and almost every user back then was a casual weekend shooter doing roughly the same kind of editing. Nobody in that room was picturing a professional retoucher who reworks 90 percent of everything by habit. Why would they. Loomstone barely had any yet.
What Dova actually did: she rebuilt the kept rate to run against each photographer's own last 90 days, not one number for the whole app, and she pulled a hand-graded set of 300 photos, scored by three senior retouchers, to set the real bar version 15 had to clear before anyone called it fixed. Checked this way, version 15 came in well outside that bar. It got pulled back for casual users within a week, patched, and re-tested against the same 300 photos before it shipped again.
The replay: four weeks after the patch, casual photographers' kept rate was back to 82 percent, close to where it started, and this time Dova had a number that would have caught the drop on day three instead of day eleven.
What I would tell myself, back before any of this: a number that keeps climbing for everyone at once was never going to warn you. You have to go build the one that can climb for one group while it falls for another, and still show you both.
SPARK, five decisions in order
This is a design question, build the measurement system before the flaw shows up, so SPARK fits. Not a diagnosis of something that already broke, and not a ranking of what to build first.
Two things worth naming straight, since this is where the real judgment sits. Loomstone looked at bolting on a one-tap rating after every edit, a simple thumbs up or down. It got turned down on purpose: forcing a rating mid edit breaks the flow the whole app is built around, and the photographers who'd actually tap it are the ones who already feel strongly, which skews the number toward complaints and skips everyone quietly satisfied. The failure worth naming by name is a signal that quietly stops meaning what it used to, the kept rate could stay perfectly healthy even while a model gets subtly worse, if the flaw is one a photographer just doesn't notice on a phone screen. The guardrail is a standing check, not a one time build: every week, a senior retoucher hand grades 40 of that week's kept photos against the golden set, checking that kept still means good, not just unnoticed. And the bar itself isn't a single clean number. A retouch model version clears the gate when its baselined kept rate stays inside a calibrated band of the version before it, across a rolling two-week window, checked against those 300 hand-graded photos, not one good day. The trade is real: baselining every photographer's own habit costs more to store and compute than one fixed rule, and waiting ten minutes after export to call an edit truly kept, instead of counting it the second it happens, keeps the dashboard a little behind the actual afternoon. That delay is the price of a number that tells the truth instead of one that's just fast.
And if you want to be sure it really works, try it somewhere else
Same five letters, a veterinary tool instead of a photo app, so the method proves itself instead of repeating a story I happened to prepare.
Briarcroft runs Notewell, an app that drafts a vet's visit notes from the audio of an exam. Devan Brenner runs product for it.
S, situation. Vets today write every note by hand after each appointment, or dictate into a recorder and type it up later that night, whichever the clinic happens to have set up.
P, payoff. Stop asking vets to rate a note. Watch whether they keep it as dictated, lightly edit it, or discard it and start over.
A, anchor. Track kept, edited, and discarded, baselined against each vet's own last 90 days of notes, not one fixed edit rate for the whole clinic.
R, risk. One emergency vet on the night shift dictates very short, clinical notes, then always adds a warm, plain-language line for the pet owner before saving, every single time, good AI draft or not. Read flat, her near-constant edit looked like a constant complaint about Notewell's writing. Baselined against her own habit, it wasn't a signal at all.
K, keep out. No attempt yet to tell a real clinical correction apart from a stylistic add. That split comes later, once the keep versus edit signal itself is trusted.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the anchor: whatever the product is, track kept, discarded, and reworked, baselined per person, never one number for everyone.
Cost: engineering says the personal baseline can't ship for two months. Don't fall back to one fixed cut off in the meantime and call it good enough. Hold the riskiest model versions back from wide release until the baseline exists.
The model got better, for real: say the retouch model gets meaningfully sharper at exposure. That's still not proof a flat cut off is safe. A better model just makes the next hidden regression harder to spot without the baseline.
Where people run it wrong.
They treat a flat topline number as proof nothing broke, instead of asking whose number is flat and whose actually moved.
They read a high edit rate as rejection across the board, instead of checking whether that person edits everything on principle.
They build the personal baseline once and never check it against real human grading again, so a signal that quietly stopped meaning anything keeps getting trusted for months.
How to use it live. Say the split before naming a single fix: "I always ask whether the behavior I'm reading means the same thing for every person doing it, because the same action, a rework, an edit, can be a complaint from one person and a habit from another." That buys you room to give the real answer instead of reaching for "add a feedback button" on reflex.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Doesn't waiting ten minutes after export before counting something as kept just let a bad model slip through in the meantime?" Response: the wait only delays counting, not detection, since it's applied evenly to every edit. A bad model still gets caught inside the same rolling two-week window, it just means each individual "kept" label is trustworthy instead of an early guess a photographer changes their mind about five minutes later.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Success metrics for AI products
- #1 What is the difference between a model metric and a product metric? Give an example of each.
- #2 Define the north star metric for an AI writing assistant and defend it.
- #3 Why is usage a weak success metric for an AI feature?
- #4 Describe three metrics that would tell you an AI feature is trusted rather than merely used.
- #5 How do you measure whether an AI feature saved users time?
- #6 What metric captures the value of an AI feature that prevents work rather than performs it?