ConceptAdvancedQuality, Cost & Token Economics / Success metrics for AI products / #14

Explain how edit distance between AI output and final user output can serve as a metric.

The direct answer
Split edit distance by what kind of edit it was, not by how many characters moved. Weight a fix to a size, a material, or a safety claim far above a fix to tone or word order, and watch that weighted share by language and product category, because it climbs weeks before returns or complaints do. A flat or falling average edit distance proves nothing on its own, since it blends a three word factual miss together with a fifty word stylistic rewrite as if they were the same kind of mistake.
Do this, in order
  1. Split edit distance by edit type, factual versus stylistic, before doing anything else with the number.Why: a raw character count treats a critical spec fix and a marketing rewrite as the same event.
  2. Weight factual edits far above stylistic ones, and watch that weighted share by language and category, not one blended average.Why: severity, not volume, is what predicts a bad listing reaching a buyer.
  3. Score the factual-edit share against a fixed golden set of listings with known-correct translations, not only against whatever sellers happen to catch.Why: sellers only fix what they notice, so seller-driven edit distance already misses the worst errors.
  4. Freeze auto-publish for a language and category once its factual-edit share crosses a set cut off for two weeks running.Why: this is the actual decision the metric has to earn, not a number that just sits on a dashboard.
  5. Leave stylistic edit share alone completely, no matter how high it runs.Why: reworded tone never hurt a buyer, so gating on it only slows sellers down for nothing.
  6. Route any frozen language-category pair to a person who checks the glossary and term list before it goes back to auto-publish.Why: the model found one wrong way to translate a term, and without a check it can find another one nobody has seen yet.

How to answer this, stage by stage

Nobody is grading whether you can define edit distance. They are grading whether you can say why a falling average can still be hiding the one mistake that matters. Seven moves get you there.

1
Scope it to one real product and one real tool
Say it like this
"Let's ground this. Sellingo runs Portside, a tool that translates a seller's product title and description into the language of wherever they're shipping to. Yevette Ostergaard runs product for it. Right now Portside handles about twenty thousand listings a day across forty languages."
Why this works
Gives the interviewer a real system to reason about, instead of leaving edit distance as an idea floating in the air.
2
Name the outcome that actually matters, before naming any number
Say it like this
"Before I name a metric, here's what has to be true. A buyer in Berlin reads that listing and understands exactly what they're about to get, no surprise when the box shows up. If Portside gets that wrong even once in a while, it doesn't matter how fast or how cheap the translation was."
Why this works
This is the L step. Naming the real outcome first stops the answer from defaulting to whatever number is already on a dashboard.
3
Say what the team was already watching, and why it looked fine
Say it like this
"The team was watching overall edit distance, how much of the AI's text a seller changed before publishing. It had been falling for months, twenty two percent down to nine percent. Everyone read that as Portside getting better, quarter over quarter."
Why this works
Naming the obvious number honestly, before taking it apart, is what makes the reversal land instead of sounding like a strawman.
4
Give the early signal, split by what changed
Say it like this
"Here's the fix. Split edit distance by what kind of edit it is, not by how many characters moved. A seller fixing a size, a material, or a safety word is a factual edit. A seller reworking tone is a style edit. Watch the factual share, by language and category, and it climbs weeks before returns do."
Why this works
This is the E step, and it matches the direct answer word for word. A vague "watch the edits closely" fails this stage.
5
Show how the raw number gets misread, both ways
Say it like this
"Raw edit distance lies twice. A seller who catches 'genuine leather' should have read 'vegan leather' fixes three words, tiny distance, huge miss. A seller who rewrites a whole sentence for tone moves fifty characters and nothing was ever wrong. Blend those into one average and the number tells you nothing about either one."
Why this works
This is the A step, the hardest one in LEAD. It moves the answer from "track the metric" to "here's exactly how the metric fools you."
6
Say what you'd actually do at the threshold
Say it like this
"Once the factual-edit share for a language and category crosses fifteen percent and stays there two weeks running, checked against a fixed set of listings with known-correct answers, that pair gets frozen from auto-publish until a person checks the glossary terms. Style edits never gate anything, no matter how high they run."
Why this works
This is the D step. It turns the metric into a decision instead of a number nobody acts on.
7
Close on the trade you're accepting, and the option you ruled out
Say it like this
"We looked at putting a person on every listing before it goes live. Ruled it out, twenty thousand listings a day, that kills the whole point of an instant translator. The trade I'd take instead: a bad listing might still slip through for a few days on any one pair, but a factual-edit spike catches the pattern before it reaches thousands of buyers, not after a complaint does."
Why this works
Naming a rejected option and a real cost is what turns "we'd monitor for that" into a decision you can defend.
If you remember one thing A number that keeps falling is not proof of anything by itself. It is only proof you have not asked what is inside it yet.

Let's learn

What happens when the one number everyone trusts to mean "getting better" is actually two very different numbers, hiding inside each other?

Portside is a tool inside Sellingo, the cross-border marketplace. A seller lists a product once in their own language, and Portside translates the title and description into whatever language the buyer's country uses, in under a minute for a batch of twenty five listings. Before Portside, most sellers paid a freelance translator about ninety cents a listing and waited two days for a batch to come back.

Yevette Ostergaard runs product for Portside. For months, the team watched one number: how much of the AI's translated text a seller changed before publishing, averaged across every language and every category. It started at twenty two percent in the first weeks after a big language launch, and it kept falling, week after week, down to nine percent by week ten. Every review meeting treated that fall as the headline win.

Twelve weeks, two lines, and only one of them was ever charted
Overall edit distance, every language blended Factual-edit share, German kids' outdoor gear only
32% 16% 0% week 0 week 10
The line everyone watched fell the whole time. The line nobody had built climbed past it around week six and kept going.

Here is the turn. The falling average was not proof of anything about the leather boots. The nine percent number was real, it just came from thousands of listings where the only edits sellers made were stylistic, "great for adventurous kids" reworded into "built for outdoor play," fifty characters changed and nothing was ever wrong. Buried inside that same average sat a much smaller group of listings where a seller fixed exactly three words, and never told anyone why.

The average was not lying about most listings. It was just standing next to the ones that mattered and calling them the same thing.

The three-word fix mattered because of what it was fixing. In a batch of rain boot listings for the German market, Portside had translated "vegan leather" as "echtes Leder," genuine leather. Not a clumsy phrasing. A confident, wrong claim about what the product was made of. A handful of sellers caught it and quietly corrected it. Most did not, because nothing about the sentence read as broken.

Knowledge spark: what is a golden set? A fixed batch of listings where a person has already written down the correct translation, on purpose, ahead of time. You score new translations against it the same way every time, instead of relying on whatever a seller happens to notice and fix.

At its worst, this cost more than a few confused buyers. Returns for that whole category had sat close to five percent for seven straight weeks, ordinary for rain boots. By week eleven, once enough of the mistranslated batch had shipped and come back, it reached fifteen percent. A reseller running three hundred and forty rain boot listings through Portside got flagged by Sellingo's own marketplace trust team for a misleading material claim. Buyers who wanted a vegan product had paid extra for one, on the strength of a word Portside itself had gotten backward. Their account was frozen for nine days while the listings were pulled and checked by hand.

Return rate, German kids' outdoor gear, before and after
16% 0% 5% 15% Weeks 1 to 7, average Week 11
This is the lagging line the factual-edit share was already predicting back at week six. Returns did not move until week nine, five weeks after the leading signal had already climbed.

The choice I would take back is not any single translation. It is that Portside scored its own health with one blended edit-distance number across every language and category, because that was the one number the whole team could compute cheaply and agree on. That was fine at launch, three language pairs, a few hundred listings a day, when a bad translation would surface fast through support tickets no matter what the dashboard said. It stopped being fine at forty languages and twenty thousand listings a day, where a handful of serious errors get buried inside a big number that still looks like it is heading the right way.

What I would leave alone: sellers reworking tone, swapping "great for adventurous kids" for "built for outdoor play," never needs a threshold or an alarm, no matter how much text moves. Nothing about a buyer's understanding of the product changes when a seller finds their own voice.

The lesson: a number that keeps falling for months can still be lying to you, if it is really two different numbers wearing one name.

Now here is the same thing as a story

The short version is above. Read on for how ordinary the week this went wrong looked from inside Sellingo.

Yevette Ostergaard has run product for Portside for two years. She wrote the first eval spec herself, sets the sample size for every model review, and can usually quote last quarter's numbers from memory before anyone opens a slide.

Early on, every two weeks, she pulled fifty real translated listings by hand and read them next to the original, regardless of what the dashboard said. She caught small things that way, a wrong unit, a stiff phrase, before they ever became a pattern. It was slow, and it was the reason she trusted the product.

As the blended edit-distance number kept falling, that habit thinned out in three quiet steps. First she cut the sample from fifty to twenty, because the number had looked healthy for three reviews running. A few months later she cut it to ten, folded into a slide she mostly skimmed. By month five she had stopped pulling samples at all. The dashboard's one line said Portside was getting better every single week, and there was always something else that needed her hour more.

Then, on an ordinary Wednesday, Tavish Priscovan from Sellingo's trust and marketplace-ops team messaged her with no warning in the subject line at all: "Can you look at this reseller's account before I escalate it."

Nobody had opened a ticket about Portside. They had opened one about a boot.

Three hundred and forty rain boot listings, translated into German, every one describing the boots as genuine leather when the source listing said vegan. Tavish had pulled the account after a spike in returns and one sharp complaint from a buyer who had specifically wanted the non-leather version and paid a premium for it. The account was frozen while trust and safety worked out how far it went.

Hand sketched comparison titled Small fix, big lie, big rewrite, no lie at all. Left panel, a listing card labelled 3 words changed with the caption vegan leather became genuine leather, nobody caught it. Right panel, a listing card labelled 50 words changed with the caption just a tone rewrite, nothing was ever wrong.
Both listings moved through Portside's review the same way. Only one of them was ever actually wrong, and it was the one that looked smallest on the ruler edit distance uses.

Yevette pulled the blended edit-distance chart first, out of habit. It still read nine percent, still falling, still the number from every recent review. She almost closed the tab. Then she pulled the raw listings instead, all three hundred and forty, and split them by what a seller had actually changed rather than by how much. About forty of the three hundred and forty had a three-word correction sitting in them: vegan swapped back in for genuine, nothing else touched. The rest had no edits at all. Every one of them still counted as a tiny edit distance, or none, inside an average that also held thousands of long, harmless stylistic rewrites from other categories that same week.

She and two engineers spent the following weekend, about thirty hours between them, manually checking four thousand recent German listings in that category for the same error before more of them reached buyers.

The choice from two years earlier came back to her clearly: the meeting where they picked one blended edit-distance number as Portside's health metric, because it was cheap to compute across every language without needing a linguist to define "factual" category by category. Splitting it felt like over-engineering for three language pairs and a few hundred listings a day. It made sense then. It stopped making sense long before anyone noticed it had.

What Yevette actually did: built the split, factual share weighted far above stylistic share, scored per language and category against a fixed golden set instead of whatever sellers happened to catch. Three weeks later, Portside launched Portuguese translations for a home goods category, and a similar single-word material swap showed up in week two. The factual-edit share for that pair crossed the threshold within eleven days. Auto-publish froze automatically, a linguist fixed the glossary entry the same afternoon, and the category was back live within forty eight hours. Zero complaints reached a buyer.

What I would tell myself, back before any of this: a number that only measures how much text moved was never going to tell you what the words meant. You have to go build the one that does, on purpose, before you need it.

LEAD, the four letters behind a number that lies while it looks healthy

This is a metric question, find the signal that moves first, so LEAD fits. Not a design question, and not a fairness question either, though the buyer paying more for the wrong material is a real cost worth naming on its own.

L
Link. The business outcome that actually matters, not the model's own score.
A buyer reading a translated listing and understanding exactly what they are about to get, with no surprise when the box arrives.
In this story: not "how fluent does the German read," but "did the buyer get what the words promised."
E
Early signal. The thing that moves weeks before the outcome does.
Factual-edit share for one language and category, tracked apart from the blended average, and weighted well above stylistic edits.
It climbed from about four percent to about twenty nine percent across ten weeks. Returns for that category did not move until week nine.
A
Abuse. How this metric gets gamed or misread.
A tiny factual fix and a huge stylistic rewrite land on the same ruler. Raw character count cannot tell severity from volume.
Forty of three hundred and forty bad listings showed a three-word edit or none at all, and still read as healthy inside the blended number.
D
Decision. What you would actually do differently at each threshold.
Freeze auto-publish for a language and category once its factual-edit share crosses fifteen percent for two weeks straight, checked against a golden set. Leave style edits alone entirely.
On the Portuguese launch, this caught the same shape of error in eleven days instead of ten weeks.

Two things worth naming directly, since this is where the AI-specific judgment actually lives. First, the alternative most teams reach for is putting a person on every listing before it publishes. That got ruled out on purpose: at twenty thousand listings a day, full human review is not a slower version of Portside, it is a different, much smaller product that happens to share a name. Second, the actual bar for a language-category pair is not "must never mistranslate a word." It is calibrated: a pair clears the gate when its factual-edit share stays under a set cut off across a rolling two-week window, checked against a golden set, not judged off one good day or one bad one. The failure mode worth naming by name is hallucination, the model stating something false with total confidence instead of just phrasing something clumsily, and the guardrail is the factual-edit gate plus the golden set, not a seller's word that something reads fine. The trade is real too: catching every factual slip before it ever ships would need full review of every listing, which brings back the two-day turnaround and the ninety-cent freelance cost that Portside exists to remove. The trade accepted instead is a short window, at most a couple of weeks on any one pair, where a bad translation can slip through before the factual-edit signal catches the pattern and a person steps in.

And if you want to be sure it really works, try it somewhere else

Same four letters, a health insurance claims team instead of a marketplace, so the method proves itself instead of repeating a story I happened to prepare.

Coverline runs Aria, a tool that drafts the explanation letter sent to a member whose claim gets denied. A claims adjuster reads the draft, edits it, and sends it. Marisol Andrade runs product for it.

L, link. A denied member reading the letter and actually understanding the real reason, in a way that holds up if they appeal, not just a letter that went out on time.
E, early signal. Factual-edit share for one claim type, how often an adjuster corrects a policy clause or a dollar figure, tracked apart from tone edits and weighted far above them.
A, abuse. An adjuster fixing one wrong dollar figure is a two-word edit, tiny distance, real error. An adjuster softening a blunt line into something gentler is a full paragraph, huge distance, nothing was ever wrong.
D, decision. Freeze auto-draft for a claim type once its factual-edit share crosses the cut off for two weeks running, checked against a set of pre-verified clause citations. Leave tone edits alone completely.

Appeal rate for pre-existing-condition denials, before and after, by cause
24% 0% 6% 6% +16 pts Before, week 0 After, week 8
The teal segment, appeals for ordinary reasons, barely moved. The whole rise came from one cause: clause citation errors an adjuster's tone-only edit never would have caught.
Same shape, different stakes At Sellingo the unwatched signal was a material word. At Coverline it is a policy clause. The E step does not change: find the number that is specific enough to move before the outcome does, and stop blending it into an average that only measures volume.

Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the split: whatever number you are tracking, separate the factual edits from the stylistic ones before you trust either.
Cost: engineering says a per-category golden set takes two months, not two weeks, to build for every language pair. Do not run every pair unmonitored in the meantime. Start with the categories carrying the highest-stakes claims, material, size, safety, price, and expand from there.
The model got better, for real: say Portside's translation model genuinely improves at fluency. That is still not the same claim as "factual accuracy improved." A better model can sound more natural while still getting one word backward with total confidence.

Where people run it wrong.
They treat a falling edit-distance average as proof the model is improving, instead of asking what kind of edits actually make up that average.
They build one factual threshold for the whole product, instead of one per language and category, so a spike in a small, high-stakes category gets diluted into nothing.
They put the fix entirely in a person's judgment, someone spot-checking listings by feel, instead of a number that gets checked every week whether anyone remembers to look or not.

How to use it live. Say the split before naming a single number: "I always ask what kind of edit is inside an edit-distance metric before I trust it, because a tiny fix and a big rewrite can carry completely opposite meanings." That buys you room to give the real answer, instead of reciting "track edit distance" on reflex.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
What framework fits a question that asks you to design a metric, and why?
Tap to flip
ANSWER
LEAD. It is a metric question, find the signal that moves first, not a design, risk, or diagnosis question.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Yevette Ostergaard, product manager for Sellingo's Portside translation tool, which turns product listings into the buyer's own language for cross-border sellers.
3 · THE LINK
What's the real outcome, the L step, this metric has to serve?
Tap to flip
ANSWER
A buyer reading a listing and understanding exactly what they are about to get, with no surprise when the box arrives. Not a fluent-sounding sentence.
4 · THE EARLY SIGNAL
What's the early signal, the E step, and what does it catch that the usual number misses?
Tap to flip
ANSWER
Factual-edit share for one language and category, weighted above stylistic edits. It climbed for weeks while the overall blended average kept falling and looking healthy.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at the time?
Tap to flip
ANSWER
Tracking one blended edit-distance number across every language and category. It made sense at three language pairs and a few hundred listings a day, when a bad translation surfaced fast through support tickets anyway.
6 · THE NUMBER
Fill in the blank: the DE kids' outdoor gear factual-edit share climbed from about 4 percent to about ___ percent before the return rate for that category ever moved.
Tap to flip
ANSWER
29 percent. The return rate for that category stayed flat for weeks after the factual-edit share had already climbed most of the way there.
7 · THE DECISION AT THRESHOLD
What's the D step, the actual decision at a threshold?
Tap to flip
ANSWER
Freeze auto-publish for a language and category once its factual-edit share crosses the cut off for two weeks running, checked against a fixed golden set. Leave stylistic edit share alone completely.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's its version of the factual-edit signal?
Tap to flip
ANSWER
Coverline's Aria, which drafts claim-denial letters. Its factual-edit signal is how often an adjuster corrects a policy clause or a dollar figure, tracked apart from tone edits, by claim type.

Check yourself Score: 0 / 0

Multiple choice
1. What is the actual early signal this answer proposes to catch a translation problem before returns do?
  • A. A lower overall edit distance average across every listing.
  • B. The factual-edit share for one language and category, weighted above stylistic edits and watched apart from the blended average.
  • C. The total number of listings translated per day.
  • D. A star rating sellers give the translation quality when they publish.
Show hint
It has to catch severity, not just volume of text changed.
Show answer
B. Splitting the metric by what kind of edit it is, and watching it per language and category, catches severity that a blended average buries.
True or false
2. True or false: Sellingo could have caught the genuine leather problem just by watching the overall blended edit distance keep falling.
  • True
  • False
Show hint
Check what the falling average was actually made of.
Show answer
False. The average kept falling the whole time. It blended thousands of harmless stylistic rewrites together with a small group of tiny, serious factual fixes, and a falling average could never show that second group on its own.
Fill in the blank
3. The DE kids' outdoor gear factual-edit share climbed from about 4 percent to about ___ percent over the ten weeks before returns for that category ever moved.
Show hint
Check the two-line chart in "Let's learn."
Show answer
29 percent. It climbed for weeks while the overall blended number was still falling and looking healthy.
Short answer
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look for the meeting memory, not a setting anyone could just turn up.
Show answer
Model answer: Tracking one blended edit-distance number across every language and category instead of splitting it by edit type. It made sense at three language pairs and a few hundred listings a day, when a bad translation would surface fast through support tickets regardless. It stopped making sense at forty languages and twenty thousand listings a day, where serious errors get buried inside a big number that still looks like it is heading the right way.
Short answer, apply it yourself
5. Think of an AI tool you use that produces a draft you then edit, an email assistant, a code assistant, anything like that. What would a factual-edit signal look like for that tool, separate from a stylistic-edit signal?
Show hint
Separate a change that fixes something wrong from a change that just fixes how it sounds.
Show answer
Model answer: A code assistant that autocompletes a function. A factual-edit signal would be how often you change actual logic after accepting a suggestion, a wrong condition, a wrong variable, a dropped safety check. A stylistic-edit signal would be renaming a variable or reformatting spacing. The two would tell very different stories even if the total number of characters changed looked the same.
True or false
6. True or false: the factual-edit threshold this answer describes should also gate sellers reworking a listing's marketing tone.
  • True
  • False
Show hint
Check "what I would leave alone" in "Let's learn."
Show answer
False. Style edits carry no risk to a buyer's understanding of the product, so gating auto-publish on them would only slow sellers down for nothing.
Before you close the answer
Why this works
Tests whether you will chase a number that is already trending the direction everyone wants to see, or ask what is actually inside it. Most candidates stop at "track edit distance" and never say what kind of edit.
Follow-up traps
"What about a language pair that doesn't have enough traffic to build a golden set for?" Response: pool it with the nearest category that shares the same kind of factual claims, material, size, safety, until it has enough volume, rather than leaving it unmonitored.

"Couldn't sellers just stop editing altogether to keep their factual-edit share low?" Response: no, because the golden set is scored independently of what sellers touch. A language pair with zero seller edits and a bad golden-set score still gets flagged.
If pressed
The golden set is not static. A quarter of its entries get swapped out every quarter, so neither the model nor a seller can quietly learn to pass one fixed batch of test listings while missing everything outside it.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more