How does metric selection change when the AI feature is free but expensive to run?
A free feature has no dollar sign standing guard over it. That does not mean there is nothing to measure. It means one number has to do the job two used to split between them.
- Track quality per dollar as the one main number, never quality alone and never cost alone.Why: chasing either one by itself breaks the feature, just from opposite directions.
- Put an automatic spend alert on the inference bill.Why: a runaway bill already has a forcing function, the invoice. Make sure someone reads it inside days, not a whole quarter.
- Run the quality half against a fixed set of test garments every week, not just once at launch.Why: nothing forces anyone to notice quality sliding on its own. A weekly check is the only alarm it will ever get.
- Default to protecting quality harder than cost, whenever neither side is flashing red.Why: the quiet quality drop takes far longer to notice and costs more trust once it's finally found. That's the real asymmetry.
- Flip to protecting cost harder once spend crosses a fixed share of the budget with no lift in paid conversion behind it.Why: past that line, the free feature stops paying for itself even indirectly.
- Flip straight back to quality the moment the weekly check drops two weeks running, whatever the bill says that month.Why: that's the earliest warning the slower, costlier failure will ever give you.
How to answer this, stage by stage
Nobody is grading whether you can say "watch both." They are grading whether you can commit to one number and defend the asymmetry underneath it. Six moves get you there.
Let's learn
What happens when a feature is free to use and expensive to run, and only one half of that sentence has a number stuck to it?
Mirrorwear is a button on Meridienne's product page. Point a phone at it, or use a stored photo, and it shows you wearing the exact jacket or dress on screen, built on an image generation model, before you ever step into a fitting room. Nothing about it charges the shopper. It never has.
The team's first success metric was a plain one: average render score, a quick one-to-five rating shoppers gave after each try-on, averaged across the week. It sat around 3.8 for the first quarter, then climbed to 4.3 over the next two, real thumbs from real shoppers. To keep it climbing, the team moved from a standard rendering pipeline at nineteen cents a render to a far more realistic one, forty six cents a render. Nobody put a cost line anywhere near that decision. The score was the only thing on the slide.
Meridienne's cloud budget system runs an automatic alert on every cost center, and it fires when one crosses one hundred ten percent of its own quarterly forecast. Nineteen days after the true fit model reached every shopper, it fired for Mirrorwear. The bill was real: on pace for about $475,000 that month, against a feature bringing in no revenue at all.
The team's fix was fast and, on its face, reasonable. Swap to a smaller, faster model at eleven cents a render, cutting weekly spend from $110,000 to under $27,000. Cancel the weekly render-score survey too, since response rates had sunk to two percent and nobody trusted it much anyway. The bill stopped being scary within a week.
Here is the turn. The cost problem was never really the problem. Watching one number in isolation was the problem, and the team had just done it a second time, in the opposite direction. Nobody was checking whether the cheaper model still rendered garments correctly, because the whole team had just spent a stressful week proving that watching a bill closely fixes a bill.
Mirrorwear had a golden set: four hundred garment photos, human graded once, covering solids, patterns, sequins, and metallics. It was only ever run at launch, for each model, never on a schedule. The true fit model had cleared it at ninety one percent. Nobody reran it when the cheap model shipped.
At its worst, this cost more than a bad render. A category manager, checking samples ahead of a patterned-heavy collection launch, found sleeves clipping through jackets and sequin patterns rendering as smeared blocks of color. It had been going on for eleven weeks. Try-on-to-purchase conversion on patterned garments had already fallen from 9.2 percent to 5.1 percent, and the collection's own launch numbers were being read as "AI try-on isn't moving the needle here," with someone in the room ready to propose cutting Mirrorwear from patterned items entirely.
The choice I would take back is not the cost ceiling. It's launching the feature with one score and no dollar figure next to it in the first place, treating quality and cost as two separate teams' problems instead of one number that has to answer for both.
What I would leave alone: Meridienne's plain-text size tip, a one-line note that suggests a size based on past orders. It costs a fraction of a cent to generate and a wrong guess costs a shopper nothing but a return label. A single satisfaction score is genuinely fine there. Quality per dollar earns its place where a mistake is expensive to make and expensive to notice, not on every AI feature in the building.
The lesson: a metric watched alone always looks healthy right up until the moment it isn't, because nothing alone can tell you when it's the wrong half of the story.
Now here is the same thing as a story
The short version is above. Read on for the Tuesday Anika almost signed off on pulling Mirrorwear from an entire collection.
Anika Sobczak has run product for Mirrorwear at Meridienne for two years, long enough to know a bad review meeting the moment a slide shows one number and nothing else.
When render score started climbing, from 3.8 toward 4.3 over two quarters, it went on every deck, and rightly so. Quality genuinely got better. Nobody had a reason yet to ask what it cost, because at launch scale it barely cost anything at all.
The habit thinned in three small steps. In the first months, Anika still asked "and what's this costing us" every time a model swap came up in a review. By the third quarter, with the score climbing and finance saying nothing, she stopped asking it in the room, not on purpose, just because nobody else was asking either. By month ten, the true fit model rollout got approved in a fifteen minute meeting on the strength of one slide: render score, 4.3, the best it had ever been.
Nineteen days later, an automated email landed in her inbox at 8:02 on a Tuesday. Not a person, a system: Mirrorwear had crossed one hundred ten percent of its quarterly infrastructure forecast, on its own, with three weeks left to go.
Anika pulled the number that had never once made a slide, cost per render, and did the math live in the follow up meeting. About $110,000 a week. She felt the floor drop, because she had approved the swap that caused it without asking the question she used to ask every time.
Callen Wexler, who runs FP&A for the product org, was on the call. He said the quiet part out loud: "We caught the bill in nineteen days. How long would it have taken us to catch a blurry sleeve?" Nobody in the room could answer him. The meeting moved on anyway, and inside ten minutes, everyone had agreed to a hard cost ceiling, because a number they could all see felt safer than a number nobody had been watching. The cheap model shipped within the week. Nobody asked the golden set to weigh in first.
Anika didn't lose sleep over the cost ceiling. She thought the problem was solved. That's the part she keeps coming back to. Nobody was hiding the quality risk that week. Nobody was even thinking about it, because the whole room had just spent an hour proving that watching a bill closely fixes a bill.
A year earlier, when Mirrorwear first shipped, the choice to track render score alone had been fast and reasonable. The feature was new, the volume was small, and cost was genuinely trivial at that scale. Nobody in that first meeting could have pictured 240,000 renders a week. The decision that made sense then just never got revisited once it stopped making sense.
The replay, run through the fixed design: one ratio, cost per render that actually clears the weekly golden set check, sits on every model-swap slide from day one. The true fit model still ships, because the extra twenty seven cents a render is worth it once it clears the bar. Six months later, when a cheaper model comes up again, it doesn't ship blind. It gets checked against the golden set first, fails at 68 percent on patterned garments, well under the 85 percent floor, and goes back for a second pass instead of reaching a single shopper. Eleven weeks later, there is no conversion drop to explain, because there is no bad render to find.
What Anika would tell herself, standing in that Tuesday meeting: the number missing from the slide is never the one everyone's already arguing about. It's the one nobody thought to put there in the first place.
PICK, for a feature that never sends the shopper a bill
This is a commit-then-defend question, not a habit with two settings, so PICK does the work here: name the pick first, then prove the asymmetry that justifies defaulting to one side over the other.
Two things worth saying out loud here, since this is exactly where an AI PM question earns its name. The alternative most candidates reach for is a flat cost ceiling with no quality check attached to it. That's exactly what Meridienne tried, and it's why the blurry sleeve took eleven weeks to catch. A ceiling protects the bill and nothing else. It can't tell you when the thing it's protecting has stopped being worth using. The real bar for shipping a cheaper model is calibrated, not a hard rule: it only reaches every shopper once it clears eighty five percent on the golden set, checked weekly, not once at launch. Nobody expects one hundred percent. An AI-generated photo is never going to nail every sequin and every pattern. The floor exists to catch the model failing on one kind of garment before a shopper does, not to demand perfection.
Name the risk plainly: this is a model quietly getting worse at one category of input, patterned and metallic fabric, while looking fine everywhere else, a distribution shift nobody was watching for because nobody had split the number by garment type. The guardrail is the golden set, run every week and broken out by category, not blended into one average a bad category can hide inside.
The trade is real too. Running that weekly check spends its own slice of inference budget, and any new model has to run for a few days on a small slice of shoppers before it reaches everyone, a few days of slower rollout in exchange for never quietly trading a shopper's trust for a lower bill.
And if you want to be sure it really works, try it somewhere else
Same four letters, an insurer instead of a retailer, so the method proves itself instead of repeating a story you happened to prepare.
Brackston sells auto insurance. Snapclaim is its free tool: a policyholder photographs their own damaged car after an accident, and a vision model estimates the repair cost on the spot, no adjuster visit needed for smaller claims. Priyank Fairweld runs the product.
P, position. Priyank picked cost per estimate that a licensed adjuster doesn't have to override, not raw estimate accuracy alone and not inference cost alone.
I, impact. If accuracy alone is tracked, finance feels a growing vision-model bill nobody flagged. If cost alone is tracked, a policyholder feels it, in a lowball estimate on their own bumper, and most don't dispute it, they just quietly shop a new insurer at renewal.
C, cost asymmetry. The claims-processing bill shows up on a monthly vendor invoice, caught inside one billing cycle, about twelve days once someone's watching. A quietly wrong estimate on hail or bumper damage has no such alarm. It surfaced at the next quarterly claims audit, fourteen weeks later, and only because override rates on one damage category had crept up enough for an analyst to ask why.
K, kill criteria. Flip to protect cost harder once vision-model spend crosses a fixed share of the claims-processing budget with no drop in manual adjuster hours to show for it. Flip back to protect quality the moment the override rate on any single damage category rises two review cycles running.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the position, name the metric, name the rejected alternative, name the asymmetry, done.
Cost: engineering says the weekly golden set check itself is eating real inference budget. Don't drop the check to save money. Narrow it to the garment categories that actually failed before, patterns and metallics, and say plainly that the trade is real: a slightly smaller check, in exchange for one that still catches the failure that actually happened.
The model got better, for real: say the rendering model's benchmark score jumps ten points overnight. That still isn't the same claim as "the weekly check no longer matters." A better model can still get quietly worse at one specific thing, patterned fabric, sequins, whatever the next blind spot turns out to be.
Where people run it wrong.
They pick whichever number is easiest to defend in this quarter's review, not the one that actually protects the product.
They treat a healthy bill as proof nothing is wrong, the same mistake as treating a healthy quality score as proof nothing is expensive.
They build the cost alert and never build the quality one, because the cost alert is the one somebody above them will actually ask about.
How to use it live. Say the position before the reasoning: "My metric is quality per dollar, not either one alone." That single sentence buys you the room to walk the asymmetry calmly, instead of listing every candidate metric you can think of and hoping the interviewer picks one for you.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if the quality drop is real but tiny, one point on one rare garment type? Isn't flipping back to protect quality overreacting?" Response: size it before acting. A one point drop on a rare category isn't the same call as a twenty three point drop on a third of the catalog, which is what actually happened here.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Success metrics for AI products
- #1 What is the difference between a model metric and a product metric? Give an example of each.
- #2 Define the north star metric for an AI writing assistant and defend it.
- #3 Why is usage a weak success metric for an AI feature?
- #4 Describe three metrics that would tell you an AI feature is trusted rather than merely used.
- #5 How do you measure whether an AI feature saved users time?
- #6 What metric captures the value of an AI feature that prevents work rather than performs it?