ConceptFoundationalModel Fluency & the AI PM Role / Working with ML engineers and researchers / #8
What information does an ML engineer need from you that a frontend engineer does not?
GUARD · what Snorri needed and never got, the week Tuyet signed a cap that wasn't there
Tesserae is Halworth AI's tool for translating a company's contracts so the other side can read and sign them. Ahtola Systems, a Finnish equipment maker, used it to localize a supplier agreement for Long Xuyen Precision, a contract manufacturer it had just started using in Vietnam. Tuyet Duong, Long Xuyen's operations director, signed the Vietnamese version in good faith. It took a dispute over defective parts, six months later, to show that the version she signed and the version Ahtola actually wrote were not saying the same thing.
The direct answer
Hand the ML engineer four things a frontend spec never has to carry: real example inputs, including the rare and messy ones, a stated bar for how often the model's allowed to be wrong and which direction of wrong is worse, a concrete way to judge whether an output is actually good, and a plan for what happens when the model isn't sure. Skip any one of those and the engineer has to guess a bar. That guess quietly becomes the real product decision, made by someone with no way of knowing they'd just made one.
Do this, in order
Hand over the four things a frontend spec skips: real edge-case examples, a stated failure tolerance, an eval spec, and a low-confidence plan.Why: skip one and the engineer invents it, and their invention becomes the actual bar nobody agreed to.
Give real, messy example inputs, not a clean demo set.Why: rare clause types and rare language pairs are exactly what a screen mockup never shows, and exactly where the model breaks first.
Name an acceptable error rate, and say which kind of error costs more.Why: without it, nobody can tell whether a 27 percent correction rate is a crisis or expected, until a customer finds out.
Write the eval spec before a line of code ships, not after a customer finds the gap.Why: "make it work" can't be checked. A reviewer agreement rate can.
Decide what happens below the point where the model isn't sure, and route it to a person instead of shipping it silently.Why: this is the one missing piece that turns an invisible model weakness into a signed legal document.
Put the four fields on the ticket template itself, not in one person's memory.Why: without a structural home, the same guess gets made again on the next rare pair.
How to answer this, stage by stage
Nobody is grading whether you can define "machine translation." They're grading whether you can turn a vague handoff gap into a specific, defensible product decision.
1
Scope it to one concrete handoff
Say it like this
"Let's make this real. Say a company builds Tesserae, a tool that translates a company's contracts so the other side can read and sign them. A product manager hands the translation model to an ML engineer the same week she'd hand a login screen to a frontend engineer. That's the gap I want to show you."
Why this works
Keeps the answer from turning into a lecture on how machine translation works.
2
Say your structure out loud
Say it like this
"I'll run this as GUARD. Groups, who's exposed if the handoff is thin. Unequal, where that harm actually lands. Ability to contest, who can push back on a gap they can't even see. Reduce, the actual fix. Detect, how you'd know it's happening before a customer does."
Why this works
Two seconds of structure tells the interviewer you have a method, not just an opinion about specs.
3
Answer the literal question, plainly, before any story
Say it like this
"A frontend engineer needs a mockup, a list of screen states, and a click flow, because the system's deterministic: same input, same output, every time. An ML engineer needs four different things. Real example inputs, including the weird ones. A stated bar for how often it's allowed to be wrong. A real way to judge if an output passed. And a plan for what happens when the model isn't sure. None of that exists for a login screen, because a login screen doesn't have a failure rate to plan for."
Why this works
The question asked for a difference. Give it in one breath before any story, or the answer never actually lands.
4
Name both people GUARD makes you name
Say it like this
"There's Snorri, the ML engineer who has to guess. And there's Tuyet Duong, an operations director in Vietnam who never sees his ticket, his pipeline, or his guess. She only ever sees the finished contract, and she signs it believing it says what it says."
Why this works
The strongest GUARD move: naming the person with no lever. Skip it and this stays an engineering footnote.
5
Reframe what's actually being tested
Say it like this
"This sounds like 'what does an ML engineer need.' It's really 'what happens when nobody defines what wrong looks like.' A frontend bug shows itself the second you click it. A wrong translation looks exactly like a right one, right up until someone downstream signs it."
Why this works
This line is the whole answer in miniature. The long version below proves it happened for real.
6
Give the committed, structural answer
Say it like this
"So here's what I'd build: a PM handoff checklist for any ML work, four required fields before an engineer can start. Real edge-case examples. A stated failure tolerance. An eval spec. A defined boundary behavior for low-confidence output. Not a wiki page nobody reads. A field on the ticket that can't stay blank."
Why this works
Ties straight back to the direct answer, and shows it's a structural fix, not a feeling.
7
Close on something checkable
Say it like this
"You'll know it's fixed when an ML engineer can point to the exact line that told them the acceptable error rate on a liability clause, instead of telling you, eight months later, that they assumed it should be as good as the pairs with more data."
Why this works
Ends on a test anyone in the room could run, not a promise that it's handled.
Let's learn
Tesserae is Halworth AI's tool for translating a company's contracts and getting them ready to sign, in whatever language the other side reads.
One list tells you what the screen looks like. The other tells you how wrong the thing behind the screen is allowed to be.
For its two biggest language pairs, Finnish to German and Finnish to English, Tesserae is genuinely good. Halworth built a real eval set for both, years ago: actual contract clauses, a bar for how many a bilingual lawyer has to agree with, a rule for what happens when the model isn't sure. Out of 812 German contracts translated last year, 12 needed a legally material fix after review, about 1.5 percent. Out of 640 English ones, 6 did, under 1 percent.
Then Ahtola Systems asked for one more pair: Finnish to Vietnamese, to localize supplier contracts for Long Xuyen Precision. Isidora Caskey, the product manager, wrote the ticket the way she'd write any new feature. Add a language to the dropdown. Reuse the existing pipeline. Ship it in the same portal. It read exactly like the ticket for a new screen state, because to her, that's what it looked like.
Knowledge spark: what's a pivot translation?
Finnish and Vietnamese don't have much text ever translated straight between them, so a model routes it through a language with more data in the middle, usually English: Finnish to English, then English to Vietnamese. Two hops instead of one. Whatever gets bent out of shape on the first hop rides straight through the second.
Snorri Bjelke, the ML engineer, built it the way the ticket described. He didn't get a legal-clause eval set for Vietnamese, because nobody had ever asked for one. He reused the general cut-off tuned on the German and English pairs, on the reasoning that a number that worked twice would probably work a third time.
Here's the turn. Over the next 14 months, Halworth shipped 22 contracts through the new pair. Not one was ever flagged for review, because nothing had been built to flag anything for a pair this small. That looked like a clean record. It wasn't a clean record. It was a pair nobody had ever told what wrong looks like.
We didn't ship a worse translator. We shipped one nobody had ever told what wrong looks like.
One of those 22 contracts was the supplier agreement between Ahtola and Long Xuyen Precision. Its liability clause, in Finnish, capped Long Xuyen's damages at the value of the order. The Vietnamese version Tuyet Duong actually read and signed dropped that cap. A scope error, picked up somewhere in the English pivot step, turned a limited promise into what read, on the page she signed, like an open one. Nobody caught it, because nothing was built to catch it. Six months later, a batch of defective parts turned into a real dispute, and Ahtola's lawyers pointed straight at the version Tuyet had signed.
The choice I would take back
Isidora wrote Tesserae's Vietnamese launch ticket exactly like a UI feature: add the language, reuse the pipeline, ship it. She never separated what a frontend engineer needs to start from what an ML engineer needs to start, because until this pair, the difference had never once cost anyone anything.
What I would leave alone: the German and English pairs don't need any of this rebuilt. They already have a real eval set, a real bar, and years of contracts behind them. This isn't a reason to slow those down. It's a reason to ask, every time a new pair ships, whether it's standing on the same ground the old ones stand on.
The lesson: a spec can be complete for one language and quietly empty for the next one, and nobody finds out by reading it. They find out when a real document goes out the door on the empty half.
Now here is the same thing as a story
The short version above is what you'd actually say in the room. Read this one for what fourteen months of a clean-looking record actually hid.
Snorri Bjelke has built translation pipelines at Halworth for five years. He is the person other engineers ask when a pair behaves strangely, because he can usually tell, from the shape of the errors alone, whether the problem is the data, the model, or the pivot step in between. When Halworth added German and English years ago, he was the one who insisted on a real eval set before either pair went live: real clauses, a bilingual lawyer's sign-off rate, a documented cut-off for what got flagged. Both pairs have run clean ever since. He built his reputation on refusing to ship a language pair on a guess.
Four questions a real ML spec answers before an engineer starts. Snorri's ticket for the Vietnamese pair answered none of them.
So when Isidora's ticket for Finnish-to-Vietnamese landed in his queue, he did what looked, from where he sat, like the same sensible thing. There was no eval set for this pair. There was no time budgeted to build one, since it was framed as a small, low-volume addition to an existing feature. So he reused the cut-off already proven on German and English, the same number that had kept two pairs clean for years, and shipped it. He wasn't being careless. He was doing the version of "checked" that the ticket in front of him made available.
For over a year, nothing pushed back. Ahtola's procurement team sent a handful of Vietnamese contracts through the self-serve portal every month, mostly routine purchase orders and delivery schedules. Snorri stopped thinking about the pair by month four. Why would he keep watching a number that had never once moved.
Isidora wrote the ticket. Snorri picked the number. Tuyet only ever saw what came out the other end.
Tuyet Duong runs operations for Long Xuyen Precision, a contract manufacturer outside Ho Chi Minh City that had just won its first order from Ahtola: machined housings, a two-year agreement, real volume. She doesn't read Finnish. She read the Vietnamese contract Halworth's portal produced, the same portal Ahtola's team used for every supplier now, and it looked exactly like every other supplier agreement she'd signed: a scope, a schedule, a price, a liability clause with numbers in it that seemed, at a glance, ordinary. She signed it on a Thursday afternoon and moved on to the next order.
Nothing about that Thursday looked like a decision that would matter. That's the part worth sitting with. Tuyet did exactly what a careful operations director does with a contract from a serious partner, in the language her own lawyers read.
Six months later, a batch of housings came back with a tooling defect Long Xuyen's own quality team had missed. Ahtola's legal team opened a claim, and instead of capping it at the order's value, the way Long Xuyen expected from every other supplier deal it had ever signed, they pointed at language in the signed Vietnamese contract that read as an open-ended liability commitment. Long Xuyen brought in outside counsel who could read both languages side by side. It took her two weeks to find it: the Finnish original capped the liability. The Vietnamese translation, the only version Tuyet had ever read, did not.
Four of these five steps worked exactly as built. The missing one, a flag on a clause the model wasn't sure about, was never built for this pair.
We did not lose a clause. We lost the one sentence that told Tuyet what she was actually agreeing to.
The decision Snorri would take back sits in a much earlier meeting, the one where the Vietnamese pair got approved as a small addition to an existing feature. Nobody in that meeting asked whether a pair with almost no direct training data needed its own bar for what counts as a pass. It felt like a reasonable question with an obvious answer: of course the same pipeline would work, it always had. Nobody had a number to argue with, because nobody had asked for one.
Run the same fourteen months again, with one change: the Vietnamese pair launches with its own real eval set, a stated tolerance for the clause types that carry legal weight, and a rule that anything below that bar routes to a bilingual reviewer instead of straight to the portal. Ahtola's routine purchase orders and delivery schedules still ship same-day, untouched, because those clauses were never the risk. The supplier agreement's liability clause, the one that sits right at the edge of what the model had ever seen for this pair, gets flagged automatically. A reviewer catches the dropped cap in eleven minutes. Tuyet signs a contract that says what the Finnish one says.
What I'd tell myself, back in that approval meeting: the shortcut wasn't wrong because Snorri was careless. It was wrong because nobody had ever asked him the one question that would have told him he needed a different number. Everyone assumed the ticket in front of him already had the answer in it. It didn't. It just had a screen.
GUARD, for a cap that vanished somewhere in the pivot
This isn't really about whether Snorri picked the right cut-off. GUARD is for naming who's exposed when a spec never says what wrong costs, and what specifically changes about the ticket.
GGroups. Who is affected, and how.
Three people, not one. Isidora Caskey, the product manager who wrote the launch ticket like a feature request. Snorri Bjelke, the ML engineer who filled the gap in that ticket with his own best guess. And Tuyet Duong, the operations director at Long Xuyen Precision, who never sees a ticket, a pipeline, or a guess, only the finished contract she's asked to sign.
Name the two who hold the spec and the one who only ever sees what comes out the other end. Most answers only name the first two.
UUnequal. Where the harm concentrates, and on whom.
It doesn't spread evenly across everything Tesserae translates. Payment terms, delivery windows, contact details, the bulk of any contract, translate the same way in every pair Halworth runs, because that language is common everywhere and the model has plenty of it to learn from. The harm concentrates in the rare clause types, liability, indemnification, warranty scope, sitting inside the rarest language pairs, exactly the combination with the least data behind it and the least attention paid to it.
This is what makes it structural, not bad luck. The clause type most likely to be mistranslated is also the one that costs the most when it is.
AAbility to contest. Who has a lever, and who has empty hands.
Snorri could have asked for an eval set and a stated tolerance, if he'd known the ticket was missing them. He didn't, because nothing about the ticket, the process, or the two pairs that came before it ever suggested this one needed something different. Tuyet had even less. She had no way to ask "does the Vietnamese version say what the Finnish one says," because nothing she was shown, the portal, the contract, the signing page, ever mentioned that a model had guessed at this, or that the guess might be wrong.
This is GUARD's sharpest move: naming who can't tell "translated" from "translated by a pipeline nobody ever told what wrong looks like," before it costs a real supplier relationship.
Four required fields, not four nice-to-haves. A ticket missing any one of these is a ticket asking the engineer to guess.
RReduce. The actual fix, not a policy memo.
Put four required fields on any ML handoff ticket, and don't let it move to "in progress" with any of them blank: real example inputs, including rare and messy ones for this specific pair; a stated tolerance for how often the model can be wrong, and which direction of wrong matters more; an eval spec, a real reviewer agreement rate, not "make it work"; and a defined boundary behavior for low-confidence output, meaning a human reviews it instead of it shipping silently.
The alternative worth naming and rejecting: pull the Vietnamese pair entirely until it has years of data behind it, the way German and English do. That protects the next Tuyet, but it also strands every routine, low-risk Vietnamese contract, the ninety percent that were never the problem, back to slow, costly outside translators. Gating only the clause types that carry legal weight behind a human review protects the one sentence that matters without shutting off the rest.
DDetect. How you'd know, before a customer does.
Track a legally-material correction rate by language pair, every quarter, not just an overall accuracy number. Over the trailing 14 months, Halworth's established pairs held steady: 1.5 percent on Finnish-German, under 1 percent on Finnish-English. The new pair never showed up on that report at all, because nobody had built a report for it. When outside counsel finally checked, during the dispute, 6 of the pair's 22 contracts needed a legally material fix, about 27 percent.
The failure worth naming plainly: a model can perform fine on average and still fail badly on the one clause type nobody ever measured separately. A dashboard that only reports the average will say everything's fine right up until it isn't.
Contracts needing a legally material fix after review, by language pair, trailing 14 months
Same portal, same underlying model family, same week. The only thing different about the third bar is that nobody had ever told this pair what wrong looks like.
Finnish-to-Vietnamese contracts shipped with zero clauses ever flagged, cumulative, by month
Cumulative contracts shipped with no clause ever flagged for review
This count climbed for over a year before anyone looked for it. Nothing on Snorri's normal dashboard would ever have shown it, because that dashboard only ever tracked the pairs that had a bar to measure against.
The test that keeps this honest
If the fix here were "tell engineers to ask more questions," nothing would actually change, that's a habit, not a design decision. The trade-off worth naming out loud: routing low-confidence clauses on new pairs to a bilingual reviewer costs real time and money. A same-day self-serve translation becomes a two or three day reviewed one, and someone has to pay a reviewer's hourly rate. That's accepted on purpose, for the clause types that carry legal weight, in exchange for a signed contract that says what the original says. It stays same-day for everything else, because everything else was never the risk.
And if you want to be sure it really works, try it somewhere else
Same five letters, a garment mill instead of a translation desk, and what's missing this time is a bar for a defect nobody had labeled yet.
Warpline, built by Loomvale Systems, watches fabric run through a mill's line and flags likely defects, streaks, slubs, tension pulls, for a technician to check before a roll ships. Osvalda Renwood owns Warpline's model pipeline the way Snorri owns Tesserae's.
Two common defects sat safely in Warpline's well-tested zone. The one that actually cost the mill money was rare enough that nobody had ever set a bar for it.
Warpline's spec, when it launched, read like a UI ticket too: detect a defect, draw a box, show a label, matching the mockups already built for an earlier quality tool. It never said which kind of miss mattered more, or what should happen below a certain confidence. Osvalda reused a general threshold from that earlier, better-documented line. Warpline was strong on common defects and silently swallowed low-confidence flags instead of surfacing them, because nobody had said what to do with those. A rare weft tension pull, the kind that shows up once every few thousand meters, sat right at that swallowed edge and got auto-dismissed for eleven straight weeks. Reinette Coppinger, the floor technician, trusted every roll Warpline didn't flag, the same way Tuyet trusted every clause Tesserae didn't flag. A retailer's cutting room found the flaw in finished garments and issued a full-batch return.
Same rank, mapped onto Warpline: give Osvalda real examples of the rare defect types, a stated tolerance for how often a rare one can slip through, an eval spec built specifically around the low-frequency defects, and a rule that anything below the confidence line gets a human look instead of a silent pass. Reinette's honest answer to a supervisor asking "did we check for this" would have taken one glance at a flag, not a full-batch recall.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: a frontend spec needs a mockup, an ML spec needs real examples, a stated error bar, an eval, and a plan for uncertainty, because only one of those two systems has a failure rate to plan for.
Cost: no budget this quarter to rebuild the whole pipeline. Gate only the clause or defect types that carry real weight first, and leave the low-risk ones running as they are; you don't need to fix everything at once to stop the one thing that actually hurts.
The model got better, for real: say a newer version of Tesserae measurably improves Vietnamese accuracy overall. Keep the review gate on the legal-weight clauses anyway. An average going up doesn't prove the one rare clause type got safer, until it's been checked on its own.
Where people run it wrong.
They treat "the model works fine most of the time" as proof it works for the one clause or defect that actually matters.
They add a disclaimer about machine translation or automated detection instead of actually defining a bar and a review gate.
They let a customer dispute or a batch return be the first real test, instead of a bilingual or expert review sample before launch.
How to use it live. Before answering, ask yourself out loud: "is this the kind of output someone might act on without double-checking it, and if so, has anyone ever told the model what wrong actually costs here?" Say which, for the specific case in the question, and the shape of the right fix usually falls right out of it.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
Which framework fits "what does an ML engineer need from you that a frontend engineer does not"?
Tap to flip
ANSWER
GUARD: groups, unequal, ability to contest, reduce, detect. It fits because the real risk isn't a definition gap, it's who gets hurt when a PM under-informs an ML engineer the same casual way they'd get away with under-informing a frontend engineer.
2 · THE PEOPLE
Who are the three people this answer names?
Tap to flip
ANSWER
Isidora Caskey, the product manager who wrote the ticket like a feature request. Snorri Bjelke, the ML engineer who filled the gap with his own guess. Tuyet Duong, Long Xuyen Precision's operations director, who signed the contract with no way to check it.
3 · THE GAP
What four things does an ML spec need that a frontend spec never has to carry?
Tap to flip
ANSWER
Real example inputs, including the rare and messy ones. A stated bar for how often the model can be wrong, and which direction costs more. A real eval spec, how a pass gets judged. A defined plan for what happens when the model isn't sure.
4 · THE UNEQUAL HARM
Where did the harm actually concentrate, and where did it never show up?
Tap to flip
ANSWER
It concentrated in rare, legally-weighted clause types (liability, indemnification) inside the rarest language pair. Routine clauses, payment terms, delivery windows, translated fine everywhere, because that language is common and well covered in every pair.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Isidora wrote the Vietnamese launch ticket exactly like a UI feature, add a language, reuse the pipeline, ship it, with no separate ML-specific fields. Nobody in the approval meeting asked whether a low-data pair needed its own bar.
6 · THE NUMBER
Fill in the blank: of Halworth's 22 Finnish-to-Vietnamese contracts, ___ needed a legally material fix, about ___ percent, versus about ___ percent on the established German pair.
Tap to flip
ANSWER
6 of 22 needed a fix, about 27 percent, versus about 1.5 percent on Finnish-German (12 of 812). Nobody found the gap until a dispute forced outside counsel to check both language versions.
7 · THE REPLAY
Same fourteen months, new ticket, what changes?
Tap to flip
ANSWER
The Vietnamese pair launches with its own eval set and a stated tolerance. The liability clause, sitting right at the edge of what the model had seen, gets flagged automatically. A bilingual reviewer catches the dropped cap in eleven minutes. Tuyet signs a contract that says what the Finnish one says.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs GUARD again on a different product. Which one, and who plays the equivalent roles?
Tap to flip
ANSWER
Warpline, Loomvale Systems' fabric-defect detector. Osvalda Renwood plays Snorri's role, owning the pipeline. Floor technician Reinette Coppinger plays Tuyet's role, trusting every roll the model didn't flag.
Check yourself Score: 0 / 0
True or false
1. True or false: the real problem in this story is that Tesserae's Vietnamese translations were less accurate than its German ones.
True
False
Show hint
Check the highlight line in Let's learn, and stage 5 of the walkthrough.
Show answer
False. Some accuracy dip on a low-data pair is expected. The real problem is that nobody defined a tolerance or an eval bar for it, so nobody could tell an acceptable dip apart from a broken clause.
Multiple choice
2. Why couldn't Snorri just reuse the cut-off tuned on the German and English pairs for the new Vietnamese pair?
A. He could have, and did. That was the right call.
B. The cut-off was tuned against a real eval set built for pairs with plenty of data. The Vietnamese pair had neither, so the same number meant something different.
C. Vietnamese is always harder for a computer to read than Finnish or English.
D. Tesserae's dropdown menu mislabeled the language pair.
Show hint
Check the knowledge spark on pivot translation, and Snorri's decision in the story.
Show answer
B. A number tuned against one eval set doesn't automatically transfer to a pair with a different amount of training data behind it, even inside the same product.
Fill in the blank
3. Of Halworth's 22 Finnish-to-Vietnamese contracts, ___ needed a legally material fix after the dispute forced a review, a rate of about ___ percent, versus about ___ percent on the Finnish-to-German pair.
Show hint
Check the bar chart in the GUARD recap.
Show answer
6, 27, and 1.5. The gap sat invisible for 14 months because nothing tracked this pair's correction rate on its own.
Short answer, name the rejected alternative
4. Besides the handoff checklist, what alternative fix does this answer name and reject, and why does it lose?
Show hint
Look at the Reduce step in the GUARD recap.
Show answer
Model answer: Pulling the Vietnamese pair entirely until it has years of data behind it, the way German and English do. It loses because it strands every routine, low-risk contract too, the ninety percent that were never the problem, forcing them back to slow, costly outside translators.
Short answer, apply it yourself
5. Think of an AI feature you've used or built where a person handed a model spec to someone else. Was there a stated bar for how wrong the model was allowed to be, or did the builder have to guess one?
Show hint
Look for anywhere a product ticket described a screen, but not a failure rate or an eval.
Show answer
Model answer: A support chatbot's escalation model, told to "flag urgent tickets," with no stated bar for how many non-urgent tickets it could wrongly flag before agents stopped trusting the flag at all, the same guessed-at gap as Snorri's, just in a lower-stakes product.
Fill in the blank, work the number
6. If Halworth had shipped 88 Finnish-to-Vietnamese contracts instead of 22, at the same 27 percent rate, roughly how many would have needed a legally material fix, and would that change the fix?
Show hint
Scale 6 out of 22 up to 88, then ask what the fix actually depends on.
Show answer
About 24. The fix doesn't change, since it depends on which clause types ever get a stated bar, not on how many contracts pass through. More volume just means the same unfixed gap costs more, faster.
Before you close the answer
Why this works
Tests whether you understand that a probabilistic system needs a different kind of spec than a deterministic one, not just different words for the same document. Most candidates can list "training data" and "accuracy" as ML-flavored nouns. Naming the exact four fields, and turning them into a structural handoff decision, is the part almost nobody does unprompted.
Follow-up traps
"Isn't this just Snorri not asking enough questions?" Response: he had no way to know there was a gap to ask about. Nothing in the ticket, the pipeline, or the two pairs that came before it ever suggested this one needed something different. The fix is structural, a required field, not a note to communicate more.
"Why not just require human review on every single contract, in every language?" Response: the German and English pairs already have years of a real eval set behind them, so review there is a cost with no matching benefit. The fix targets the actual gap, the pairs and clause types with no proven bar, not translation as a whole.
If pressed
The Finnish clause used a common construction, "vastuu rajoittuu X:aan," where the limiting word sits at the very end of the sentence. The English pivot step flattened that into a passive phrase and let the limit attach to the wrong part of the sentence. The Vietnamese output inherited that flattened version. It's a known failure shape in pivot translation for languages whose clauses put the limiting word in different places, not a random glitch.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.