Explain the difference between a defect and an acceptable error rate to a non-technical executive.
Listwell drafts the property description an agent posts to the county MLS, pulling from the listing sheet and the agent's own photos. Marrowick Realty Group, 340 agents across three counties, built it in-house. Isolina Marchmont owns its product story. Thaddeus Corrigan, Marrowick's Designated Broker, is the one whose real estate license sits behind every word an agent publishes under the Marrowick name. Fourteen months in, a cost-driven model swap taught them both that a healthy accuracy score and a real defect can share the exact same dashboard.
- Split Listwell's fields into two budgets: cosmetic, where 3 in 100 is fine, and seven critical facts, where it isn't.Why: mixing them lets a healthy blended score quietly cover for a defect hiding in the tiny group that can get someone sued.
- Watch the critical-field error rate on its own weekly line, never folded into the overall accuracy score.Why: seven fields are a sliver of the roughly 38 Listwell touches per listing, so a real spike in just those seven barely dents the blended average.
- Set a real cutoff, 3 per 1,000 listings, and require two straight weekly checks before acting.Why: one bad week on a small sample is noise. Two running is a trend worth a rollback.
- Auto-rollback the model and page Thaddeus the moment the cutoff is crossed twice, don't wait for a complaint to notice.Why: the Oakmire listing had already gone public before anyone had looked at the split number at all.
- Never let a healthy blended score justify keeping a critical-field regression live.Why: that's the exact sentence, "97 percent, still inside the band," that explained away four weeks of a real defect.
- Leave the cosmetic-field budget alone.Why: tightening it blanket-wide makes every listing blander without shrinking the actual risk sitting in the seven fields.
How to answer this, stage by stage
Nobody is grading whether you can say "not all mistakes are equal." They're grading whether you can split them into two real groups, each with a real number attached.
Let's learn
Here's what a healthy score can hide. A dashboard can sit at 97 percent for fourteen months straight while seven of its fields quietly cross into something that isn't a quality problem anymore. It's a defect.
Listwell drafts the property description an agent posts to the county MLS, pulling from the listing sheet and the agent's own photos. Marrowick Realty Group built it in-house for the 340 agents working across three counties.
Before Listwell, writing a listing description by hand took an agent about 35 minutes on average. Across the brokerage, that's roughly 580 hours a month, every month, just on the writing.
With Listwell, a draft appears in under a minute. An agent reads it, fixes a word or two, and posts it, about 4 minutes of review. Combined across the brokerage, that's about 67 hours a month instead of 580.
Listwell touches about 38 fields per listing. Checked against the source records, it's been right 97.4 percent of the time for fourteen straight months, a 2.6 percent error rate. Almost all of it is cosmetic: a generic sentence, a word choice an agent swaps out in seconds.
Seven of those 38 fields are different. Bedroom count. Bathroom count. Square footage. Lot size. HOA dues. Whether there's a pool. The assigned elementary school. Get one of those wrong and it isn't a style note, it's a fact a buyer, a regulator, or a lawsuit can hold Marrowick to. Baseline, Listwell got one of these wrong about once every two months, across the whole brokerage.
Three months ago Marrowick absorbed a smaller brokerage, adding about 150 agents and pushing Listwell's monthly draft volume toward 1,500. To keep the AI bill from doubling with it, the platform team swapped in a smaller, faster model, cutting the cost of a single draft from 44 cents to 16. The blended accuracy score barely moved, 97.4 down to 97.0 by week four. The critical-field rate did something else entirely.
The extra mistakes weren't the real problem. Seven fields drifting from about one in 2,000 listings to nearly four in 1,000, in a single month, while the number everyone was watching stayed inside the band it had always lived in, that was the real problem.
What it costs at its worst: a wrong school zone goes out, a buyer's agent flags it to the county Realtor association because school-boundary claims are specifically watched for steering, and the listing gets pulled mid multiple-offer window. Thaddeus, whose own license sits behind every word an agent publishes, wants Listwell switched off company-wide, for everyone, over a mistake confined to seven of 38 fields.
What I would leave alone: the cosmetic fields. A 3 percent miss rate on tone and phrasing costs nothing, an agent fixes it in the same four minutes they were already spending on review. Tightening that budget wouldn't touch the real risk. It would just make every listing read a little blander.
The lesson: "acceptable error rate" was never one number for Listwell to hit. It's a question you ask field by field. Some mistakes get a budget, because being wrong there costs nothing. Some get a number close enough to zero that two bad weeks pulls the plug, because being wrong there costs a license.
Now here is the same thing as a story
Read the short version above when you're in the room. Read this one when you want to feel why a single house on Oakmire Lane, not a bad model, is what put Thaddeus's license on the table.
Esme Rundgren has sold houses in the same three counties for eleven years. Hand her a listing sheet and she can tell you, before she's finished her coffee, which line the buyers will ask about first.
Listwell arrived fourteen months ago. For most of that first year, Esme read every draft against the county sheet before she touched publish, bedroom count, square footage, the school zone, all of it. Every single time, it checked out. She started skimming the fields she'd already checked twice by month four, and by month eight, she wasn't opening the county sheet at all anymore. Why check work that always checked out.
In week three of the new model, Listwell drafted a listing for a house on Oakmire Lane. The draft read, "just minutes from top-rated Oakmire Elementary." The parcel's actual assigned school, three miles further, is Tanner Hollow Elementary, rated well below Oakmire on the state's own index. Esme read the draft once, it sounded right, and she published it.
A buyer's agent caught the mismatch two days later and flagged it to the county Realtor association's advertising line, since school-boundary claims are exactly the kind of thing that line watches for. The listing had to come down mid multiple-offer window. One of the three interested buyers walked before it went back up. Marrowick's compliance officer filed a correction. And Esme, mortified, went back to opening the county sheet on every single field, for every single listing, the way she had fourteen months ago.
Three days after that, Thaddeus called Isolina into his office with the county's letter in his hand. He wanted Listwell off, company-wide, until someone could promise him it would never happen again.
The decision Isolina would take back traced to a short meeting the week Listwell first launched. Someone had asked whether the acceptable-error-rate policy, 3 percent, needed a separate, tighter number for a handful of fields. The room decided one policy was simpler. At the time, every field failed about as rarely as every other one, so a split budget looked like process for its own sake.
Run the same week three again, with the split rule already running. The critical-field rate crosses 3 per 1,000 that Monday. Engineering gets 48 hours, no page yet, just a look. It crosses again the following Monday, 3.9 per 1,000, and the rule fires on its own: the model rolls back to the version from before the merger, the seven fields get a second confirmation step, and Thaddeus gets a same-day page, not a letter from the county three weeks later.
One version of that month ends with a public complaint and a demand to kill the whole tool. The other ends with an engineering ticket nobody outside the team ever hears about.
What Isolina would tell herself, back in that first launch meeting: a single 3 percent budget wasn't simpler. It just meant nobody had decided, in advance, which seven mistakes the brokerage would never agree to make.
LEAD: the seven fields with no acceptable error rate at all
Not a way to argue a low score is actually fine. LEAD forces you to say, out loud, which mistakes were priced in on purpose and which ones never were.
And if you want to be sure it really works, try it somewhere else
Same four letters, a hospital radiology department instead of a brokerage, and this time a wrong-side finding, not a wrong school, is the mistake with no acceptable rate at all.
Radlign drafts the narrative section of a radiology report from a radiologist's structured findings, at Kestrahaven Regional Health. Renata Okafor owns its reporting quality. Every radiologist signs every report before it ships, unlike Marrowick's agents, so the parallel isn't about a missing check. It's about what a busy radiologist rubber-stamps when a draft reads exactly the way a normal one should.
Same rank, different lever, mapped straight onto LEAD. The link is the same shape: not Radlign's 91 percent acceptance rate, but whether a wrong-side finding reaches a signed report, since a laterality error can point a surgeon at the wrong kidney. The early signal is the same shape too: the laterality-and-measurement error rate, checked weekly on its own line, never folded into the blended 91 percent. The abuse is identical: "acceptance is still at 90.6, basically unchanged" used to wave off a handful of laterality slips hiding inside it. And the decision runs the same tiers: cross 2 per 1,000 reports twice running, and the model rolls back the same day, with a mandatory second radiologist confirming laterality and measurement until it clears.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: split the fields that can get someone hurt or sued from the ones that can't, give the first group a number close to zero, and act the moment it's crossed twice.
Cost: no budget to keep the expensive model everywhere. Keep the cheap one for the cosmetic fields. Route only the handful of critical fields through a stricter, slightly slower check.
The model got better, for real: say Listwell's next model genuinely improves across the board. The split doesn't go away. A better average can still be built entirely on the 31 fields it already handled well.
Where people run it wrong.
They watch one blended number and assume a small, expensive slice can't be failing underneath a healthy average.
They tighten the acceptable rate for everything at once, which mostly makes the safe fields worse without fixing the risky ones.
They wait for a public complaint to tell them what a split weekly number would have told them a month earlier.
How to use it live. Ask the split question before naming a number: "Which facts in this draft would cost real money or real trust if even one of them were wrong, separate from the ones that just read a little flat." That question alone usually finds the seven fields before you've named a single percentage.
One thing worth naming directly, since this is where the real judgment sits. Priya Sandhu's platform team considered tightening Listwell's whole acceptable-error-rate policy from 3 percent to 1 percent across all 38 fields, rather than splitting the critical ones out. It lost, because most of that 2.6 percent lives in cosmetic wording an agent fixes for free, so a blanket cut would have made every listing blander without shrinking the actual risk sitting in seven fields. The AI-specific failure worth naming is grounding, not raw capability: the cheaper model could recognize Oakmire Elementary as a real, nearby, well-rated school, it just wasn't reliably tying that name to the parcel's actual assigned-school code in the structured record, especially on a field that barely varies in the training data a model like it would have seen. The guardrail is a field-level grounding check that cross-verifies the seven critical fields against the structured source record before publish, regardless of which base model drafted the prose, backed by the weekly critical-field line as ongoing production monitoring. And the trade-off is real and accepted on purpose: the grounding check adds about two seconds to every draft, only on those seven fields, so Marrowick keeps the cheaper model's savings everywhere it's safe to and pays a small, deliberate latency cost only where being wrong is expensive.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Why not just tighten the whole 3 percent acceptable rate to be safe?" Response: most of that 3 percent lives in phrasing an agent fixes in seconds. Tightening it everywhere makes every listing blander without shrinking the risk sitting in seven fields.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on What changes when the product is probabilistic
- #1 Name three product decisions that change when a feature's output is probabilistic rather than deterministic.
- #2 A traditional feature either works or has a bug. Explain why that framing breaks for an LLM feature.
- #3 What does 'correct' mean for a summarization feature? Give a definition your engineering team could test against.
- #4 QA files a bug that reads: the model gave a wrong answer once. How do you triage it?
- #6 Why can you not write an acceptance criterion like 'the output must be accurate' for a generative feature?
- #7 Describe how you would set a quality bar for a feature whose output is free text.