ConceptAdvancedQuality, Cost & Token Economics / Success metrics for AI products / #24

How does metric selection change when the AI feature is free but expensive to run?

A free feature has no dollar sign standing guard over it. That does not mean there is nothing to measure. It means one number has to do the job two used to split between them.

The direct answer
When there is no revenue number to lean on, the main metric has to be quality per dollar, not quality by itself and not cost by itself. Chasing quality alone runs up a bill nobody is watching. Chasing cost alone lets the feature quietly get worse until people stop coming back and never say why. Protect the quality side harder by default, because it has no invoice to catch it. Only flip that default when the spend itself starts costing more than the feature is worth.
Do this, in order
  1. Track quality per dollar as the one main number, never quality alone and never cost alone.Why: chasing either one by itself breaks the feature, just from opposite directions.
  2. Put an automatic spend alert on the inference bill.Why: a runaway bill already has a forcing function, the invoice. Make sure someone reads it inside days, not a whole quarter.
  3. Run the quality half against a fixed set of test garments every week, not just once at launch.Why: nothing forces anyone to notice quality sliding on its own. A weekly check is the only alarm it will ever get.
  4. Default to protecting quality harder than cost, whenever neither side is flashing red.Why: the quiet quality drop takes far longer to notice and costs more trust once it's finally found. That's the real asymmetry.
  5. Flip to protecting cost harder once spend crosses a fixed share of the budget with no lift in paid conversion behind it.Why: past that line, the free feature stops paying for itself even indirectly.
  6. Flip straight back to quality the moment the weekly check drops two weeks running, whatever the bill says that month.Why: that's the earliest warning the slower, costlier failure will ever give you.

How to answer this, stage by stage

Nobody is grading whether you can say "watch both." They are grading whether you can commit to one number and defend the asymmetry underneath it. Six moves get you there.

1
Scope it to one real product, one real number
Say it like this
"Let's ground this. Meridienne is an online fashion retailer. Mirrorwear is the free button on a product page that shows you wearing the jacket you're looking at, using an image model. Anika Sobczak runs it. It costs real money every time someone presses it, and it has never charged a shopper a cent."
Why this works
Grounds the tradeoff in a real product and a real number before naming a single metric.
2
Say your structure out loud
Say it like this
"Here's how I'd take this apart. Pick a side first, then say who gets hurt by each kind of wrong metric, then find the real asymmetry between them, then say what would ever make me switch."
Why this works
Two seconds of structure tells the interviewer you have a plan before you name a single metric.
3
Commit to a position before any reasoning
Say it like this
"My pick is quality per dollar, not quality alone and not cost alone. Concretely: cost per render that actually clears a weekly fidelity check against a fixed set of test garments."
Why this works
This is the direct answer. An interviewer testing a tradeoff wants the pick first, not five candidate metrics.
4
Name who feels each kind of error
Say it like this
"If I only watch quality, finance feels it first, in a bill that grows for months before anyone looks. If I only watch cost, the shopper feels it first, in a jacket that renders with a warped sleeve, and she just closes the tab. Neither team sees the other one's problem coming."
Why this works
Names both sides in the person's terms, not the metric's, which is the whole point of the tradeoff step.
5
Find the real asymmetry and say which side you protect by default
Say it like this
"Here's the asymmetry. The bill has a forcing function, the invoice shows up whether anyone asked for it or not. Quality decay has no forcing function at all, nobody's job is to notice a render looking a little worse this week than last. So by default I protect quality harder, because it's the side that can go wrong in silence for months."
Why this works
This is the line the whole tradeoff turns on. Skip it and you've just picked a metric, not defended one.
6
Give the kill criteria, then close on one line
Say it like this
"I'd flip and protect cost harder the day spend crosses a fixed share of the cloud budget with nothing to show for it in paid conversion. And I'd flip straight back to quality the moment the weekly check drops two weeks running, no matter how good the bill looks that month. Either way it's one ratio, checked every week, never two dashboards that never talk to each other."
Why this works
Kill criteria are what separate a confident pick from a stubborn one.
If you remember one thing The side with no invoice is the side that needs the metric most. A bill announces itself. A quietly worse render does not.

Let's learn

What happens when a feature is free to use and expensive to run, and only one half of that sentence has a number stuck to it?

Mirrorwear is a button on Meridienne's product page. Point a phone at it, or use a stored photo, and it shows you wearing the exact jacket or dress on screen, built on an image generation model, before you ever step into a fitting room. Nothing about it charges the shopper. It never has.

The team's first success metric was a plain one: average render score, a quick one-to-five rating shoppers gave after each try-on, averaged across the week. It sat around 3.8 for the first quarter, then climbed to 4.3 over the next two, real thumbs from real shoppers. To keep it climbing, the team moved from a standard rendering pipeline at nineteen cents a render to a far more realistic one, forty six cents a render. Nobody put a cost line anywhere near that decision. The score was the only thing on the slide.

Cost per render, before and after chasing a higher quality score
50c 25c 0 Standard model True fit model 19c 46c
At 240,000 renders a week, the switch to the higher fidelity model turned about $45,600 a week into just over $110,000 a week, close to $475,000 a month, for a feature that has never sold a single subscription.
We did not spend more to make the feature better. We spent more and never told anyone what better was supposed to buy.
The decision that mattered Launching Mirrorwear with one score, render quality, and no dollar figure sitting anywhere near it. Quality and cost lived on two different teams' slides.

Meridienne's cloud budget system runs an automatic alert on every cost center, and it fires when one crosses one hundred ten percent of its own quarterly forecast. Nineteen days after the true fit model reached every shopper, it fired for Mirrorwear. The bill was real: on pace for about $475,000 that month, against a feature bringing in no revenue at all.

The team's fix was fast and, on its face, reasonable. Swap to a smaller, faster model at eleven cents a render, cutting weekly spend from $110,000 to under $27,000. Cancel the weekly render-score survey too, since response rates had sunk to two percent and nobody trusted it much anyway. The bill stopped being scary within a week.

Here is the turn. The cost problem was never really the problem. Watching one number in isolation was the problem, and the team had just done it a second time, in the opposite direction. Nobody was checking whether the cheaper model still rendered garments correctly, because the whole team had just spent a stressful week proving that watching a bill closely fixes a bill.

Knowledge spark: what's a golden set? A fixed pile of test photos with a known right answer, checked by a person once, then reused every time a new model version ships. It's how you catch a model quietly getting worse at one kind of input without redoing the human check by hand every single week.

Mirrorwear had a golden set: four hundred garment photos, human graded once, covering solids, patterns, sequins, and metallics. It was only ever run at launch, for each model, never on a schedule. The true fit model had cleared it at ninety one percent. Nobody reran it when the cheap model shipped.

Garment fidelity score, eleven weeks after the fast model shipped, reconstructed from render logs
95% 75% 55% 85% floor week 0 week 11, caught here
Fidelity on patterned and metallic garments, about a third of the catalog, slid from 91 percent to 68 percent, well under the 85 percent floor, with no weekly check in place to catch it crossing the line.

At its worst, this cost more than a bad render. A category manager, checking samples ahead of a patterned-heavy collection launch, found sleeves clipping through jackets and sequin patterns rendering as smeared blocks of color. It had been going on for eleven weeks. Try-on-to-purchase conversion on patterned garments had already fallen from 9.2 percent to 5.1 percent, and the collection's own launch numbers were being read as "AI try-on isn't moving the needle here," with someone in the room ready to propose cutting Mirrorwear from patterned items entirely.

The choice I would take back is not the cost ceiling. It's launching the feature with one score and no dollar figure next to it in the first place, treating quality and cost as two separate teams' problems instead of one number that has to answer for both.

What I would leave alone: Meridienne's plain-text size tip, a one-line note that suggests a size based on past orders. It costs a fraction of a cent to generate and a wrong guess costs a shopper nothing but a return label. A single satisfaction score is genuinely fine there. Quality per dollar earns its place where a mistake is expensive to make and expensive to notice, not on every AI feature in the building.

The lesson: a metric watched alone always looks healthy right up until the moment it isn't, because nothing alone can tell you when it's the wrong half of the story.

Now here is the same thing as a story

The short version is above. Read on for the Tuesday Anika almost signed off on pulling Mirrorwear from an entire collection.

Anika Sobczak has run product for Mirrorwear at Meridienne for two years, long enough to know a bad review meeting the moment a slide shows one number and nothing else.

When render score started climbing, from 3.8 toward 4.3 over two quarters, it went on every deck, and rightly so. Quality genuinely got better. Nobody had a reason yet to ask what it cost, because at launch scale it barely cost anything at all.

The habit thinned in three small steps. In the first months, Anika still asked "and what's this costing us" every time a model swap came up in a review. By the third quarter, with the score climbing and finance saying nothing, she stopped asking it in the room, not on purpose, just because nobody else was asking either. By month ten, the true fit model rollout got approved in a fifteen minute meeting on the strength of one slide: render score, 4.3, the best it had ever been.

Nineteen days later, an automated email landed in her inbox at 8:02 on a Tuesday. Not a person, a system: Mirrorwear had crossed one hundred ten percent of its quarterly infrastructure forecast, on its own, with three weeks left to go.

Hand sketched comparison diagram titled Which one costs more. Left, a small plain box labeled The bill, caught in 19 days. Right, a larger jagged red orange box labeled The blurry sleeve, caught in 11 weeks.
One side had an alarm built in from day one. The other side had never had one at all.

Anika pulled the number that had never once made a slide, cost per render, and did the math live in the follow up meeting. About $110,000 a week. She felt the floor drop, because she had approved the swap that caused it without asking the question she used to ask every time.

It was never really about the four hundred and seventy five thousand dollars. It was that nobody in that room, her included, had thought to ask what number should sit next to render score before anyone signed off.

Callen Wexler, who runs FP&A for the product org, was on the call. He said the quiet part out loud: "We caught the bill in nineteen days. How long would it have taken us to catch a blurry sleeve?" Nobody in the room could answer him. The meeting moved on anyway, and inside ten minutes, everyone had agreed to a hard cost ceiling, because a number they could all see felt safer than a number nobody had been watching. The cheap model shipped within the week. Nobody asked the golden set to weigh in first.

Anika didn't lose sleep over the cost ceiling. She thought the problem was solved. That's the part she keeps coming back to. Nobody was hiding the quality risk that week. Nobody was even thinking about it, because the whole room had just spent an hour proving that watching a bill closely fixes a bill.

A year earlier, when Mirrorwear first shipped, the choice to track render score alone had been fast and reasonable. The feature was new, the volume was small, and cost was genuinely trivial at that scale. Nobody in that first meeting could have pictured 240,000 renders a week. The decision that made sense then just never got revisited once it stopped making sense.

The replay, run through the fixed design: one ratio, cost per render that actually clears the weekly golden set check, sits on every model-swap slide from day one. The true fit model still ships, because the extra twenty seven cents a render is worth it once it clears the bar. Six months later, when a cheaper model comes up again, it doesn't ship blind. It gets checked against the golden set first, fails at 68 percent on patterned garments, well under the 85 percent floor, and goes back for a second pass instead of reaching a single shopper. Eleven weeks later, there is no conversion drop to explain, because there is no bad render to find.

What Anika would tell herself, standing in that Tuesday meeting: the number missing from the slide is never the one everyone's already arguing about. It's the one nobody thought to put there in the first place.

PICK, for a feature that never sends the shopper a bill

This is a commit-then-defend question, not a habit with two settings, so PICK does the work here: name the pick first, then prove the asymmetry that justifies defaulting to one side over the other.

P
Position. State the pick before any reasoning.
Cost per render that actually clears a weekly fidelity check against a fixed, four hundred photo golden set. Not render score alone. Not cost per render alone.
I
Impact. Who feels each kind of error.
If quality alone is tracked, finance feels it first, in a bill that grows quietly for months. If cost alone is tracked, the shopper feels it first, in a bad render she never reports, she just leaves and doesn't try the next garment.
C
Cost asymmetry. Which error is cheap and visible, which is hidden and expensive.
The bill has a forcing function built in, the invoice, caught in nineteen days once a ceiling exists. Quiet quality decay has no forcing function at all. It took eleven weeks and a manual check before anyone even started looking.
K
Kill criteria. What evidence flips the pick.
Flip to protect cost harder once spend crosses a fixed share of the cloud budget for two months straight with no lift in paid conversion. Flip back to quality the moment the weekly check drops two weeks running, whatever the bill says.

Two things worth saying out loud here, since this is exactly where an AI PM question earns its name. The alternative most candidates reach for is a flat cost ceiling with no quality check attached to it. That's exactly what Meridienne tried, and it's why the blurry sleeve took eleven weeks to catch. A ceiling protects the bill and nothing else. It can't tell you when the thing it's protecting has stopped being worth using. The real bar for shipping a cheaper model is calibrated, not a hard rule: it only reaches every shopper once it clears eighty five percent on the golden set, checked weekly, not once at launch. Nobody expects one hundred percent. An AI-generated photo is never going to nail every sequin and every pattern. The floor exists to catch the model failing on one kind of garment before a shopper does, not to demand perfection.

Name the risk plainly: this is a model quietly getting worse at one category of input, patterned and metallic fabric, while looking fine everywhere else, a distribution shift nobody was watching for because nobody had split the number by garment type. The guardrail is the golden set, run every week and broken out by category, not blended into one average a bad category can hide inside.

The trade is real too. Running that weekly check spends its own slice of inference budget, and any new model has to run for a few days on a small slice of shoppers before it reaches everyone, a few days of slower rollout in exchange for never quietly trading a shopper's trust for a lower bill.

And if you want to be sure it really works, try it somewhere else

Same four letters, an insurer instead of a retailer, so the method proves itself instead of repeating a story you happened to prepare.

Brackston sells auto insurance. Snapclaim is its free tool: a policyholder photographs their own damaged car after an accident, and a vision model estimates the repair cost on the spot, no adjuster visit needed for smaller claims. Priyank Fairweld runs the product.

P, position. Priyank picked cost per estimate that a licensed adjuster doesn't have to override, not raw estimate accuracy alone and not inference cost alone.
I, impact. If accuracy alone is tracked, finance feels a growing vision-model bill nobody flagged. If cost alone is tracked, a policyholder feels it, in a lowball estimate on their own bumper, and most don't dispute it, they just quietly shop a new insurer at renewal.
C, cost asymmetry. The claims-processing bill shows up on a monthly vendor invoice, caught inside one billing cycle, about twelve days once someone's watching. A quietly wrong estimate on hail or bumper damage has no such alarm. It surfaced at the next quarterly claims audit, fourteen weeks later, and only because override rates on one damage category had crept up enough for an analyst to ask why.
K, kill criteria. Flip to protect cost harder once vision-model spend crosses a fixed share of the claims-processing budget with no drop in manual adjuster hours to show for it. Flip back to protect quality the moment the override rate on any single damage category rises two review cycles running.

Same shape, different stakes At Meridienne, the hidden cost was a shopper's trust in a jacket photo. At Brackston, it's a driver's trust in what their own insurer says their car is worth. The asymmetry holds either way: whatever can drift for a whole quarter unnoticed outranks whatever would embarrass itself inside two weeks.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the position, name the metric, name the rejected alternative, name the asymmetry, done.
Cost: engineering says the weekly golden set check itself is eating real inference budget. Don't drop the check to save money. Narrow it to the garment categories that actually failed before, patterns and metallics, and say plainly that the trade is real: a slightly smaller check, in exchange for one that still catches the failure that actually happened.
The model got better, for real: say the rendering model's benchmark score jumps ten points overnight. That still isn't the same claim as "the weekly check no longer matters." A better model can still get quietly worse at one specific thing, patterned fabric, sequins, whatever the next blind spot turns out to be.

Where people run it wrong.
They pick whichever number is easiest to defend in this quarter's review, not the one that actually protects the product.
They treat a healthy bill as proof nothing is wrong, the same mistake as treating a healthy quality score as proof nothing is expensive.
They build the cost alert and never build the quality one, because the cost alert is the one somebody above them will actually ask about.

How to use it live. Say the position before the reasoning: "My metric is quality per dollar, not either one alone." That single sentence buys you the room to walk the asymmetry calmly, instead of listing every candidate metric you can think of and hoping the interviewer picks one for you.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What framework fits this question, and why?
Tap to flip
ANSWER
PICK. It's a tradeoff between two single-number extremes, quality alone versus cost alone, and the job is to commit to a combined pick, then prove the asymmetry that justifies it.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Anika Sobczak, who runs product for Mirrorwear at Meridienne, and Callen Wexler, who runs FP&A for the product org.
3 · THE POSITION
What is the picked metric, defined exactly?
Tap to flip
ANSWER
Cost per render that clears a weekly fidelity check against a fixed, four hundred photo golden set. Not render score alone. Not cost per render alone.
4 · THE ASYMMETRY
Which error is cheap and visible, and which is hidden and expensive?
Tap to flip
ANSWER
A runaway inference bill is cheap and visible, caught in 19 days once a spend ceiling exists. Silent quality decay is hidden and expensive, it took 11 weeks and a manual check.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at the time?
Tap to flip
ANSWER
Launching Mirrorwear with render score as the only metric, no cost figure anywhere near it. It made sense when volume was small and cost was trivial, and stopped making sense once volume hit 240,000 renders a week.
6 · THE NUMBER
Fill in the blank: garment fidelity on patterned items fell from 91 percent to ___ percent over 11 weeks, while try-on-to-purchase conversion on those items fell from 9.2 percent to ___ percent.
Tap to flip
ANSWER
68 percent; 5.1 percent. Both numbers moved for eleven weeks with nobody watching either one.
7 · THE REPLAY
Same bad Tuesday, new design, what changes?
Tap to flip
ANSWER
The cheap model gets checked against the golden set before rollout, fails at 68 percent on patterned garments, well under the 85 percent floor, and goes back for a second pass instead of shipping to every shopper. No blurry sleeve ever reaches a customer.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's its version of the picked metric?
Tap to flip
ANSWER
Snapclaim, Brackston's free photo damage assessment tool. Its picked metric is cost per estimate that a licensed adjuster doesn't have to override.

Check yourself Score: 0 / 0

Fill in the blank
1. The main metric this answer picks is quality per ___.
Show hint
Check the direct answer at the top of the page.
Show answer
Dollar. Quality alone hides a runaway bill, and cost alone hides a quietly worse feature. Only a combined ratio catches both.
Multiple choice
2. If Meridienne only tracks render quality, with no cost figure next to it, who feels the resulting mistake first?
  • A. The shopper, right away.
  • B. Finance, months later, in a bill nobody flagged.
  • C. The model vendor, immediately.
  • D. Nobody. A quality-only metric has no downside.
Show hint
Think about which team is watching the number that's missing.
Show answer
B. A quality-only metric can climb for months while the bill behind it grows unwatched, until someone in finance finally looks.
True or false
3. True or false: in this story, the silent quality decay was harder to catch than the runaway inference bill.
  • True
  • False
Show hint
Compare how each one got caught, not how big either number was.
Show answer
True. The bill had a built-in alarm, the automatic spend alert, and got caught in 19 days. The quality drop had no alarm at all and took 11 weeks and a manual check.
Short answer
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look for the choice made when Mirrorwear was brand new, not a dial anyone could just turn back up.
Show answer
Model answer: Launching Mirrorwear with render score as the only success metric, no cost figure sitting anywhere near it. It made sense when volume and cost were both small, and stopped making sense once volume reached 240,000 renders a week and nobody had revisited the choice.
Short answer, apply it yourself
5. Think of a free feature you've used that must cost real money to run behind the scenes. What's one number the team could be watching alone that would hide the opposite kind of failure?
Show hint
Pick one side, quality or cost, and ask what it can't see.
Show answer
Model answer: A free photo background remover in a mobile app. If the team only watches a satisfaction score, a runaway server bill for high resolution exports can grow for months unseen. If they only watch server cost, a quietly worse cutout, jagged edges around hair or fingers, drives people away before anyone connects it to a cheaper model swap.
Fill in the blank
6. The automatic alert that caught Mirrorwear's runaway bill fires when a cost center crosses ___ percent of its own quarterly forecast, and it caught this one in ___ days.
Show hint
Check the paragraph right after the first chart in "Let's learn."
Show answer
110 percent; 19 days. That threshold and that speed only exist because the alert was built ahead of time. Quality had no equivalent, which is exactly the gap the picked metric is meant to close.
Before you close the answer
Why this works
Tests whether you'll build one honest number instead of two separate ones that never get read side by side. Most candidates either forget cost matters at all, or panic and only watch cost the moment someone above them notices the bill.
Follow-up traps
"Isn't checking a golden set every week itself expensive? Doesn't that fight the whole point of controlling cost?" Response: yes, and that's the accepted trade. A few cents of eval inference and a short trial run on a small slice of shoppers before a cheaper model reaches everyone, in exchange for never shipping a model that's already failing on a garment type nobody's looking at.

"What if the quality drop is real but tiny, one point on one rare garment type? Isn't flipping back to protect quality overreacting?" Response: size it before acting. A one point drop on a rare category isn't the same call as a twenty three point drop on a third of the catalog, which is what actually happened here.
If pressed
The golden set itself isn't static either. Meridienne refreshes it every quarter with new photos pulled from that season's actual arrivals, ten percent of the set at a time, because a set built entirely from last year's catalog stops testing the fabrics and cuts that are actually shipping right now.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more