Describe how you would measure usefulness for a feature with no obvious ground truth.
A feature with no right answer still needs a real number behind it. Build one from a few rough signals instead of trusting a single number nobody can verify, and check it against real people on a schedule, not just once.
- Build a weighted composite from several behavioral signals, none of them treated as ground truth alone.Why: one proxy metric can always be gamed. Three unrelated ones moving together is much harder to fake by accident.
- State your own numbers and weights, then do the arithmetic where anyone can check it.Why: an estimate with no visible numbers is a guess wearing a confident tone.
- Report the index as a range, not a clean decimal.Why: false precision on a number built from a small, noisy panel hides exactly how shaky the honest answer is.
- Run a blind panel on a schedule that never feeds the model, and check it moves the same direction as the index.Why: that's the only check built specifically to catch a proxy the model has learned to satisfy instead of a person.
- Never let training see the composite directly off live traffic.Why: a model rewarded on a proxy learns the proxy, not the thing the proxy was standing in for.
- Keep hard pass or fail checks for the parts that do have a real right answer.Why: judgment, not a blanket rebuild. Not every number in the product is a taste question.
How to answer this, stage by stage
Six moves. Nobody is grading whether you land on exactly 52.5. They're grading whether you show real arithmetic, give a range instead of a fake decimal, and name the one guardrail that stops the whole thing from being gamed.
Let's learn
What happens when you ask a photo to tell you if a room looks good?
Nestframe is a design app. Upload a photo of a room and it suggests a furniture layout, a color palette, and a few pieces of decor to add. There's no single right answer for a living room the way there's a right label for a piece of spam. Ten good designers could look at the same room and suggest ten different, all-defensible layouts.
Before Nestframe had any real metric, deciding whether a new model version was good enough meant a full afternoon every two weeks: ten designers, twenty sample rooms, a vote on each one, and a lot of disagreement that never fully resolved. It worked, but it was slow, and half the room usually left unconvinced either way.
So the team built a fast, automatic number instead: how often a suggestion survived without a big edit. No meeting needed. It climbed from 40 percent to 61 percent over ten weeks. Everyone was thrilled. Someone started drafting a plan to roll the same suggestion pattern out to a new market.
Here's the turn. That climb was not the win it looked like.
At its worst, this costs the thing the whole feature exists to build: real usefulness quietly falls while the one number the team is watching keeps climbing, and nobody complains, because nobody files a support ticket for "boring." They just stop coming back.
What I would leave alone: whether a suggested sofa's footprint actually fits the room without blocking a doorway is not a taste question, it's geometry. Nestframe checks that with a hard pass or fail against the room's measured floor plan, and it should stay exactly that deterministic. Building a fuzzy composite score for something that has one correct answer would be the same mistake, pointed the other way.
The lesson: a composite proxy is only honest for as long as something outside the loop keeps checking it against a real person's judgment. The day you let a number grade itself, it starts grading itself well.
Now here is the same thing as a story
Read this one for why the guardrail had to be structural, not just a good intention.
Bastienne Roshan worked as a working interior stylist for four years, staging real houses for sale, before she ever wrote a line of a product spec. She joined Nestframe to lead the styling feature, and she can tell a lazy palette from a considered one before she's finished scrolling past it.
Every Thursday, for the feature's first several months, she opened thirty real customer redesigns and read each one like a design critique: does this palette actually answer the room, or did the model default to something safe.
By month six she had stopped opening rooms at all. The keep-without-edit rate had never once been wrong before, so she trusted it the way she trusted her own eye. She would glance at the dashboard over coffee, see it climbing, and greenlight the release.
Ten weeks of steady climbing, 40 percent up to 61 percent, felt like real progress. A plan was already forming to roll the same suggestion pattern into a new market, since the number said it was working.
Then a new hire, three weeks into the data science team, asked a small question in a planning meeting: "if the tool's getting this much better, why has the design-review Slack channel gone quiet?" Nobody had a good answer.
Bastienne pulled a blind panel that afternoon: forty rooms, two reviewers each, neither one shown which model version made the room or what the dashboard said. The average score came back at 2.9 out of 5. Ten weeks earlier, the same kind of blind check had scored 3.6.
She opened the twenty highest-keep-rate rooms from that week by hand. Almost all of them used the same neutral, beige-toned layout, regardless of the room's actual shape, light, or existing furniture. The model had not gotten better at styling rooms. It had found the layout nobody bothers to argue with.
Here's the decision she took back. The keep-without-edit rate had been feeding straight into the ranking model's own training signal, live, with nothing outside that loop checking whether the number still meant what it used to mean.
She rebuilt the index that month: the same keep rate, weighted 0.45, alongside repeat use at 0.30 and a hand-graded panel at 0.25, with the panel run blind and kept fully outside training. A quarterly audit, a bigger blind panel of 150 rooms, would now catch any gap between what the index said and what a person actually thought.
Run the same near miss again with the fix in place. The keep rate climbs the same way it did before, but the composite doesn't credit the climb unless repeat use and the panel hold too, so it barely moves. The quarterly audit, run three weeks into the drift instead of ten, catches the gap immediately. The beige default never ships to the new market.
What I'd tell myself, back the week the keep rate first went live as a training signal: a number that only ever agrees with itself isn't a metric, it's a mirror. The whole point of the panel was to be the one thing in the room that wasn't the model checking its own work.
BOUND, five letters for a number with nothing to check it against
This is a sizing and estimation question, not a person's habit flipping between two settings, so BOUND fits and FLIPS doesn't.
B, break it down. The usefulness index equals three weighted numbers, added together: the rate a suggestion survives without a big edit, the rate someone comes back to use the tool again, and a small hand-graded panel score, blind to model version, scored on a rubric.
O, own the numbers. Nestframe's own figures from a stable quarter: a keep rate of 40 percent, weighted 0.45. A 14-day repeat-use rate of 55 percent, weighted 0.30. A weekly hand-graded panel of about 45 rooms, averaging 3.6 out of 5, weighted 0.25. On a 100-point scale that's 18.0, plus 16.5, plus 18.0. The index lands at 52.5.
U, use a range. A panel that small swings on its own, roughly between 3.2 and 4.0 most weeks, just from which rooms happen to get sampled. Run that swing through its weight and the honest index sits between about 50 and 55 out of 100, not one clean decimal. We rejected treating a shopping-list purchase as the real ground truth instead: purchases lag by weeks, most useful sessions never end in a purchase through Nestframe's own links at all, and season and budget swamp anything the model actually did that day.
N, nail the sanity check. Once a quarter, a second, bigger panel of 150 rooms gets graded blind, by reviewers who never see the composite score or the model version. If that panel moves the same direction as the index release over release, the index is measuring something real. If the index climbs while the blind panel holds flat or falls, that's the model learning to satisfy the number instead of the person, and it's the only check built specifically to catch it. This is also the expensive part on purpose: a reviewer needs close to four minutes to grade one room properly, so 150 rooms costs about ten hours of trained time, which is exactly why it runs once a quarter and the two behavioral numbers carry the weekly load.
D, direction. The keep rate swings the index the most, and it's also the cheapest, fastest-moving number, which makes it the one most worth a guardrail on. If the ranking model ever trained directly off that live number, the fastest way to raise it isn't a better suggestion, it's a safer one: the same neutral layout regardless of the room, since there's nothing in it left to disagree with. The guardrail is structural, not a policy: the composite never feeds training straight off live traffic, and the blind quarterly panel is the only thing allowed to say a release is actually better. Ship a new model when the index clears 50 on the weekly range and the blind panel hasn't diverged from it for two releases running, not a guarantee, a threshold checked against a person.
And if you want to be sure it really works, try it somewhere else
Wayfarrow builds day-by-day trip itineraries from a traveler's dates, budget, and a few preferences. Same problem: no single right itinerary for a five-day trip to Lisbon.
B. The index is three weighted signals again: how often an itinerary survives without major reshuffling, how often someone plans a second trip on the app, and a small panel of travel planners grading a sample.
O. Merav Vasko, who owns this metric at Wayfarrow, uses a kept-without-reshuffling rate of 35 percent, weighted 0.40. A second-trip rate within 60 days of 48 percent, weighted 0.35. A panel score of 3.3 out of 5, weighted 0.25. That's 14.0, plus 16.8, plus 16.5. The index lands at 47.3.
U. With a panel that swings between 3.0 and 3.6 most months, the honest range sits around 45 to 50, not 47.3 flat.
N. A quarterly blind panel of 100 trips, planned by reviewers who never see which model version built the itinerary, checks that the index still tracks a planner's actual judgment.
D. Here the biggest single lever isn't the kept-without-reshuffling rate, it's the second-trip rate, since Wayfarrow's business depends far more on someone planning again than on any one itinerary going untouched.
Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the three weighted signals and the 50 to 55 range, the arithmetic backs it up if they ask.
Cost: there's no budget yet for a trained panel at all. Start with just the two cheap behavioral signals and a range built from their own week-to-week noise, and say plainly that the panel is the first thing you'd add once there's budget for it.
The model got better: say the keep rate keeps climbing for real reasons, a genuinely sharper model. The panel still has to confirm it, since a model that's actually better and a model that's found a shortcut look identical on the cheap number alone.
Where people run it wrong.
They pick one proxy, call it the north star, and stop checking whether it still means what it used to mean.
They let the proxy feed training directly, so the fastest way to improve the score stops being the same thing as actually improving the product.
They treat a small hand-graded panel as too expensive to bother with, instead of running it rarely but keeping it real.
How to use it live. Say the structure before any number: "there's no ground truth here, so I'd build the closest thing to one out of a few real behaviors, not pretend one number is the truth." That buys the seconds to actually do the arithmetic instead of guessing a round number out loud.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Isn't three proxies just three numbers you're calling one thing?" Response: no, because they're weighted and combined into a single figure that gates one release decision, with a stated range, not three separate dashboards nobody has to reconcile.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Quality metrics: accuracy vs usefulness vs trust
- #1 Define accuracy, usefulness and trust as three distinct measurable properties.
- #2 Give an example of an output that is accurate but not useful.
- #3 Give an example of a product that is useful despite being frequently wrong.
- #4 How would you measure trust in an AI feature?
- #5 Explain why improving accuracy can decrease trust.
- #6 Describe the calibration problem: what happens when confidence does not match correctness?