Explain how edit distance between AI output and final user output can serve as a metric.
- Split edit distance by edit type, factual versus stylistic, before doing anything else with the number.Why: a raw character count treats a critical spec fix and a marketing rewrite as the same event.
- Weight factual edits far above stylistic ones, and watch that weighted share by language and category, not one blended average.Why: severity, not volume, is what predicts a bad listing reaching a buyer.
- Score the factual-edit share against a fixed golden set of listings with known-correct translations, not only against whatever sellers happen to catch.Why: sellers only fix what they notice, so seller-driven edit distance already misses the worst errors.
- Freeze auto-publish for a language and category once its factual-edit share crosses a set cut off for two weeks running.Why: this is the actual decision the metric has to earn, not a number that just sits on a dashboard.
- Leave stylistic edit share alone completely, no matter how high it runs.Why: reworded tone never hurt a buyer, so gating on it only slows sellers down for nothing.
- Route any frozen language-category pair to a person who checks the glossary and term list before it goes back to auto-publish.Why: the model found one wrong way to translate a term, and without a check it can find another one nobody has seen yet.
How to answer this, stage by stage
Nobody is grading whether you can define edit distance. They are grading whether you can say why a falling average can still be hiding the one mistake that matters. Seven moves get you there.
Let's learn
What happens when the one number everyone trusts to mean "getting better" is actually two very different numbers, hiding inside each other?
Portside is a tool inside Sellingo, the cross-border marketplace. A seller lists a product once in their own language, and Portside translates the title and description into whatever language the buyer's country uses, in under a minute for a batch of twenty five listings. Before Portside, most sellers paid a freelance translator about ninety cents a listing and waited two days for a batch to come back.
Yevette Ostergaard runs product for Portside. For months, the team watched one number: how much of the AI's translated text a seller changed before publishing, averaged across every language and every category. It started at twenty two percent in the first weeks after a big language launch, and it kept falling, week after week, down to nine percent by week ten. Every review meeting treated that fall as the headline win.
Here is the turn. The falling average was not proof of anything about the leather boots. The nine percent number was real, it just came from thousands of listings where the only edits sellers made were stylistic, "great for adventurous kids" reworded into "built for outdoor play," fifty characters changed and nothing was ever wrong. Buried inside that same average sat a much smaller group of listings where a seller fixed exactly three words, and never told anyone why.
The three-word fix mattered because of what it was fixing. In a batch of rain boot listings for the German market, Portside had translated "vegan leather" as "echtes Leder," genuine leather. Not a clumsy phrasing. A confident, wrong claim about what the product was made of. A handful of sellers caught it and quietly corrected it. Most did not, because nothing about the sentence read as broken.
At its worst, this cost more than a few confused buyers. Returns for that whole category had sat close to five percent for seven straight weeks, ordinary for rain boots. By week eleven, once enough of the mistranslated batch had shipped and come back, it reached fifteen percent. A reseller running three hundred and forty rain boot listings through Portside got flagged by Sellingo's own marketplace trust team for a misleading material claim. Buyers who wanted a vegan product had paid extra for one, on the strength of a word Portside itself had gotten backward. Their account was frozen for nine days while the listings were pulled and checked by hand.
The choice I would take back is not any single translation. It is that Portside scored its own health with one blended edit-distance number across every language and category, because that was the one number the whole team could compute cheaply and agree on. That was fine at launch, three language pairs, a few hundred listings a day, when a bad translation would surface fast through support tickets no matter what the dashboard said. It stopped being fine at forty languages and twenty thousand listings a day, where a handful of serious errors get buried inside a big number that still looks like it is heading the right way.
What I would leave alone: sellers reworking tone, swapping "great for adventurous kids" for "built for outdoor play," never needs a threshold or an alarm, no matter how much text moves. Nothing about a buyer's understanding of the product changes when a seller finds their own voice.
The lesson: a number that keeps falling for months can still be lying to you, if it is really two different numbers wearing one name.
Now here is the same thing as a story
The short version is above. Read on for how ordinary the week this went wrong looked from inside Sellingo.
Yevette Ostergaard has run product for Portside for two years. She wrote the first eval spec herself, sets the sample size for every model review, and can usually quote last quarter's numbers from memory before anyone opens a slide.
Early on, every two weeks, she pulled fifty real translated listings by hand and read them next to the original, regardless of what the dashboard said. She caught small things that way, a wrong unit, a stiff phrase, before they ever became a pattern. It was slow, and it was the reason she trusted the product.
As the blended edit-distance number kept falling, that habit thinned out in three quiet steps. First she cut the sample from fifty to twenty, because the number had looked healthy for three reviews running. A few months later she cut it to ten, folded into a slide she mostly skimmed. By month five she had stopped pulling samples at all. The dashboard's one line said Portside was getting better every single week, and there was always something else that needed her hour more.
Then, on an ordinary Wednesday, Tavish Priscovan from Sellingo's trust and marketplace-ops team messaged her with no warning in the subject line at all: "Can you look at this reseller's account before I escalate it."
Three hundred and forty rain boot listings, translated into German, every one describing the boots as genuine leather when the source listing said vegan. Tavish had pulled the account after a spike in returns and one sharp complaint from a buyer who had specifically wanted the non-leather version and paid a premium for it. The account was frozen while trust and safety worked out how far it went.
Yevette pulled the blended edit-distance chart first, out of habit. It still read nine percent, still falling, still the number from every recent review. She almost closed the tab. Then she pulled the raw listings instead, all three hundred and forty, and split them by what a seller had actually changed rather than by how much. About forty of the three hundred and forty had a three-word correction sitting in them: vegan swapped back in for genuine, nothing else touched. The rest had no edits at all. Every one of them still counted as a tiny edit distance, or none, inside an average that also held thousands of long, harmless stylistic rewrites from other categories that same week.
She and two engineers spent the following weekend, about thirty hours between them, manually checking four thousand recent German listings in that category for the same error before more of them reached buyers.
The choice from two years earlier came back to her clearly: the meeting where they picked one blended edit-distance number as Portside's health metric, because it was cheap to compute across every language without needing a linguist to define "factual" category by category. Splitting it felt like over-engineering for three language pairs and a few hundred listings a day. It made sense then. It stopped making sense long before anyone noticed it had.
What Yevette actually did: built the split, factual share weighted far above stylistic share, scored per language and category against a fixed golden set instead of whatever sellers happened to catch. Three weeks later, Portside launched Portuguese translations for a home goods category, and a similar single-word material swap showed up in week two. The factual-edit share for that pair crossed the threshold within eleven days. Auto-publish froze automatically, a linguist fixed the glossary entry the same afternoon, and the category was back live within forty eight hours. Zero complaints reached a buyer.
What I would tell myself, back before any of this: a number that only measures how much text moved was never going to tell you what the words meant. You have to go build the one that does, on purpose, before you need it.
LEAD, the four letters behind a number that lies while it looks healthy
This is a metric question, find the signal that moves first, so LEAD fits. Not a design question, and not a fairness question either, though the buyer paying more for the wrong material is a real cost worth naming on its own.
Two things worth naming directly, since this is where the AI-specific judgment actually lives. First, the alternative most teams reach for is putting a person on every listing before it publishes. That got ruled out on purpose: at twenty thousand listings a day, full human review is not a slower version of Portside, it is a different, much smaller product that happens to share a name. Second, the actual bar for a language-category pair is not "must never mistranslate a word." It is calibrated: a pair clears the gate when its factual-edit share stays under a set cut off across a rolling two-week window, checked against a golden set, not judged off one good day or one bad one. The failure mode worth naming by name is hallucination, the model stating something false with total confidence instead of just phrasing something clumsily, and the guardrail is the factual-edit gate plus the golden set, not a seller's word that something reads fine. The trade is real too: catching every factual slip before it ever ships would need full review of every listing, which brings back the two-day turnaround and the ninety-cent freelance cost that Portside exists to remove. The trade accepted instead is a short window, at most a couple of weeks on any one pair, where a bad translation can slip through before the factual-edit signal catches the pattern and a person steps in.
And if you want to be sure it really works, try it somewhere else
Same four letters, a health insurance claims team instead of a marketplace, so the method proves itself instead of repeating a story I happened to prepare.
Coverline runs Aria, a tool that drafts the explanation letter sent to a member whose claim gets denied. A claims adjuster reads the draft, edits it, and sends it. Marisol Andrade runs product for it.
L, link. A denied member reading the letter and actually understanding the real reason, in a way that holds up if they appeal, not just a letter that went out on time.
E, early signal. Factual-edit share for one claim type, how often an adjuster corrects a policy clause or a dollar figure, tracked apart from tone edits and weighted far above them.
A, abuse. An adjuster fixing one wrong dollar figure is a two-word edit, tiny distance, real error. An adjuster softening a blunt line into something gentler is a full paragraph, huge distance, nothing was ever wrong.
D, decision. Freeze auto-draft for a claim type once its factual-edit share crosses the cut off for two weeks running, checked against a set of pre-verified clause citations. Leave tone edits alone completely.
Swap the trigger and it still runs.
Speed: an interviewer caps the answer at ninety seconds. Skip straight to the split: whatever number you are tracking, separate the factual edits from the stylistic ones before you trust either.
Cost: engineering says a per-category golden set takes two months, not two weeks, to build for every language pair. Do not run every pair unmonitored in the meantime. Start with the categories carrying the highest-stakes claims, material, size, safety, price, and expand from there.
The model got better, for real: say Portside's translation model genuinely improves at fluency. That is still not the same claim as "factual accuracy improved." A better model can sound more natural while still getting one word backward with total confidence.
Where people run it wrong.
They treat a falling edit-distance average as proof the model is improving, instead of asking what kind of edits actually make up that average.
They build one factual threshold for the whole product, instead of one per language and category, so a spike in a small, high-stakes category gets diluted into nothing.
They put the fix entirely in a person's judgment, someone spot-checking listings by feel, instead of a number that gets checked every week whether anyone remembers to look or not.
How to use it live. Say the split before naming a single number: "I always ask what kind of edit is inside an edit-distance metric before I trust it, because a tiny fix and a big rewrite can carry completely opposite meanings." That buys you room to give the real answer, instead of reciting "track edit distance" on reflex.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Couldn't sellers just stop editing altogether to keep their factual-edit share low?" Response: no, because the golden set is scored independently of what sellers touch. A language pair with zero seller edits and a bad golden-set score still gets flagged.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Success metrics for AI products
- #1 What is the difference between a model metric and a product metric? Give an example of each.
- #2 Define the north star metric for an AI writing assistant and defend it.
- #3 Why is usage a weak success metric for an AI feature?
- #4 Describe three metrics that would tell you an AI feature is trusted rather than merely used.
- #5 How do you measure whether an AI feature saved users time?
- #6 What metric captures the value of an AI feature that prevents work rather than performs it?