Why does a probabilistic product need a feedback mechanism that a deterministic one does not?
Frondwise is Hollowquill's plant identification engine. Photograph a plant and it names the species, flags whether it's safe around pets, and hands back a watering and light plan. Ottoline Marsh owns its quality bar. Ten weeks after Frondwise opened in the Gulf Coast, a cat named Basil chewed the fern the app had called pet safe, and Ottoline found out the golden set behind that 98 percent confidence score had exactly four regional photos to learn from.
- Build a feedback loop before launch, not after a scare.Why: a probabilistic model has no fixed rate you test once and trust forever, only a real rate you have to keep watching.
- Put a one-tap "get this wrong?" flag right on every result.Why: a report path buried in a settings menu collects almost nothing, so the loop never actually fills.
- Track the error rate by region and species, not one blended accuracy number.Why: a regional look-alike can be wrong constantly while the national average still reads like 96 percent.
- Re-run the eval set every time the underlying model version changes, not only at the original launch.Why: a routine vendor update can quietly shift a look-alike call, and nothing in the product announces that it happened.
- Add a hard-coded pet-safety check that runs after the model, no matter its confidence score.Why: a floor the model's own guess can't fully promise, for the one mistake nobody can undo once a pet's already eaten the plant.
- Watch the correction rate weekly against a stated bar, don't wait for a support ticket.Why: a ticket only shows what someone happened to notice, not how often the same mistake is really happening.
How to answer this, stage by stage
Nobody's grading whether you can say "AI needs more testing." They're grading whether you know a deterministic feature earns trust once and a probabilistic one has to keep earning it, and whether you can say what actually replaces the one-time test.
Let's learn
How sure does an app have to sound before nobody thinks to double check it?
Frondwise is Hollowquill's plant identification engine. Point a phone at a leaf, a stem, a pot somebody left on a porch, and it names the species, flags whether it's safe around pets, and hands back a watering and light plan, all in about twelve seconds.
Before Frondwise, a gardener staring at a plant nobody remembered buying spent close to 20 minutes cross-referencing forums, an old plant book, maybe two other ID apps, before feeling sure enough to actually act on the answer.
By its fifth month, Frondwise had about 210,000 households using it. Its launch checklist required one thing before ship: a benchmark run against a golden set of confirmed photos. It scored 96 percent overall, the best number the model had ever posted.
That same month, Hollowquill opened Frondwise to the Gulf Coast and the rest of the Southeast, about 38,000 new households over the following ten weeks. Roughly 24 percent of them, about 9,100 households, had a pet listed on their profile. Fern-family photos alone came in at about 85 a week from the new region.
Here's the turn. The 96 percent was never really the number that mattered, because a blended average like that can hide one narrow slice being wrong far more often than the rest, and nobody had a way to see that slice on its own.
What it costs at its worst: Frondwise could be quietly telling a real slice of its pet-owning households that a mildly toxic look-alike is safe, in a region where the model had barely been tested, and nobody would know, because a great launch score reads like a fact and not like a measurement of one moment.
What I would leave alone: once Frondwise has a species locked in, the watering and light schedule it hands back is a fixed lookup table, water every six days, indirect light, same answer every time for the same fern. That part isn't a fresh guess, so it doesn't drift, and it doesn't need a feedback loop. Only the identification step, the part that's actually guessing, does.
The lesson: a score measured once tells you the model was right on the day you tested it, on the photos you happened to include. It was never a promise about next Tuesday, and treating it like one is the whole mistake.
Now here is the same thing as a story
Read the short version above when you're in the room. Read this one when you want to feel why a 96 percent score, checked once, quietly stopped being a fact the day nobody was looking anymore.
Ottoline Marsh spent three years at Hollowquill before Frondwise existed as an app, running QA on the older, deterministic parts of the gardening site: the price calculator, the delivery-zone checker, things that either worked or they didn't. She also kept forty pots on her own back porch and could name half of them by leaf shape alone. When Frondwise moved from idea to build, she was the obvious person to own its quality bar.
For the six months of beta, Ottoline ran a review every single week. Fifty recent identifications, pulled at random, sent to three consulting horticulturists to mark right or wrong. Early on, that review caught real problems, a mislabeled succulent here, a watering plan that was off by half. Each one got fixed. Each week's review came back a little cleaner than the last.
The habit thinned in three beats, and every one of them felt sensible at the time. Around month three, with review after review coming back nearly spotless, she cut the sample from fifty to twenty. Around month five, she moved from weekly to monthly, because nothing had moved in ages and three horticulturists cost real money to keep booking. The week before launch, she didn't schedule a review at all. The model was about to get its real test, she told herself, against the full golden set, once, properly.
The golden set came back at 96 percent, the highest score Frondwise had ever posted, well past the 90 percent bar Hollowquill required to ship. Ottoline forwarded it to her director with one line: "Ship it, this is as sure as we're going to get." She meant it. Nothing in three years of deterministic launches had ever taught her that a number checked once could quietly stop being true.
Frondwise launched. The Gulf Coast and the rest of the Southeast opened four months later, 38,000 new households in ten weeks, fresh nursery partnerships, early reviews that read like the beta all over again. In week three, Frondwise's vendor pushed a routine update to the underlying vision model, broadening what it could recognize nationally. Nobody re-ran the golden set against it. The launch checklist had never said to.
In week ten, Rosaleen Cressmoor photographed a hanging fern she'd bought at a coastal nursery. Frondwise came back fast and certain: Boston fern, 98 percent match, safe for pets, water weekly, indirect light. She hung it in her sunroom, right where her cat Basil could reach it. Basil chewed a frond that evening. A little drooling, a little vomiting, a worried trip to the emergency vet, and a diagnosis that had nothing to do with anything dangerous: it wasn't a Boston fern at all. It was an asparagus fern, a real look-alike, mildly toxic to cats, common on Gulf Coast porches and almost absent from the two founding regions Frondwise had originally trained on.
Basil was fine within a day. Rosaleen filed one annoyed support ticket. Bevan Oyelaran, the engineer who traces Frondwise's misses, found the exact generation within the hour, saw the model had leaned on the fern's overall shape instead of the leaflet pattern that actually tells the two apart, patched a rule to weight that pattern higher, and closed the ticket.
Ottoline went to write the same closing note she'd written dozens of times on the deterministic features, and stopped halfway through the sentence. This wasn't a lookup-table typo with one wrong row. Nothing in the code was broken. The model had simply weighed one visual cue over another, on this one generation, and nothing said that couldn't happen again tomorrow, to a household with nobody watching the screen before checkout.
What Ottoline did next wasn't ask Bevan to be more careful. She pulled 240 fern-family requests logged from the region over those same ten weeks and checked each photo against Frondwise's answer by hand, three evenings at her kitchen table. Eight more were asparagus fern, called Boston fern, safe for pets. None had ever been reported. Eight in 240, about 1 in 30.
For six months, Ottoline had a rhythm. Then one great score, checked a single time, took its place, and she never noticed the day the rhythm actually ended, because reading a great number doesn't feel like closing a door.
The decision she'd take back sat in a much smaller meeting, the one where Frondwise's launch checklist got written down. It was copied almost line for line from Hollowquill's older feature launches: pass one benchmark, ship, done. Nobody in that meeting added a line for what happens after ship, or for what happens when the model's version quietly changes six weeks later, because at the time Ottoline's weekly beta review was still running, and the gap didn't feel like a gap yet.
Run the same week ten again, with the fix in place. Frondwise still occasionally weighs shape over leaflet pattern, because that's what a probabilistic model does sometimes. But now a one-tap flag sits on every result, and four Gulf Coast households flag the exact same fern mix-up within six days of the model update reaching them. The weekly dashboard, split by region and species, shows fern-family accuracy in the Southeast drop below its bar before a fifth household ever sees the wrong answer. And the week-three model version bump would have triggered a mandatory re-run of an expanded golden set before it ever went live at all.
One design let a great score stand in for ten weeks of nobody looking. The other looks every week, whether the score is great or not.
What Ottoline would tell herself, the day she typed "ship it, this is as sure as we're going to get" under a number she'd only measured once: that number was never a fact. It was a photograph of one Tuesday, and she filed it away like it would never age.
FLIPS, or why "it passed" stops being an answer
Not a story wearing a framework's clothes. This is what to actually run, in order, any time a question asks why a probabilistic feature needs something a deterministic one never did.
Two things worth naming directly, since this is where the real judgment sits. Hollowquill's team actually considered a cheaper fix first: a hard denylist, so "asparagus fern" could never be shown as safe for pets. That got rejected, because the failure was never the model naming a banned species. It was the model naming the wrong species entirely, so the denylist would sit there, unused, since the app never rendered the word "asparagus fern" in the first place. The AI-specific failure worth naming by name is confident wrongness riding on a silent model change: Frondwise displayed a 98 percent match on the wrong species, three weeks after a routine vendor update to the underlying vision model that nobody re-tested against the region's look-alikes. The guardrail is a deterministic check that runs after the model, not instead of it: any species carrying a toxicity flag gets held for a second confirmation before the app ever says "safe for pets," no matter how confident the model looked, and any model version change triggers a mandatory re-run of an expanded, region-aware golden set before it goes live. That check costs something real, a fraction of a second on every plant ID, and real review hours to keep the golden set current, in exchange for a floor under the one mistake a family can't take back once a cat's already eaten the plant.
And if you want to be sure it really works, try it somewhere else
Same five letters, a furnace instead of a fern, and this time the flip isn't about who stopped watching. It's about what a technician quietly starts doing to the input before the model ever hears it.
Wrenchnote is Marrowvent's field tool: a technician describes what they're seeing and hearing at a furnace or an AC unit out loud, and Wrenchnote drafts a written diagnostic report plus a suggested repair, replacing a paper checklist a tech used to fill out by hand. Ignatius Cray owns its accuracy rubric.
F Halcyon Petrescu, a senior technician who's called correctly on the fussiest units in her district for eleven years. L She stopped describing a call the messy way she actually thinks it through, "compressor sounds off, might be the capacitor, might not, hard to tell over the fan," because Wrenchnote's report reads cleaner once she tidies that into one confident line first. I A different flip entirely, the pre-editing flip: she feeds Wrenchnote the smoothed version of what she found, not the real one, because a messy dictation used to trip the model into a low-confidence, unhelpful report, and a clean one never does. P Wrenchnote's rubric was built entirely from tidy pilot dictations, recorded by trained pilot users describing textbook cases in a quiet office, so the accuracy score never had a way to tell "the model handles ambiguity well" apart from "technicians stopped giving it any ambiguity to handle." S The replay: Wrenchnote adds a flag for "it missed something I said," tracks how short and clean a dictation runs as its own signal, and pulls a weekly sample of real, un-tidied field audio into review. Halcyon's next uncertain call gets logged as uncertain and held for a second look, instead of reading as confident and wrong.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: build the correction flag and the region-and-species dashboard before ship, and re-test every model version change, not just the first one.
Cost: there's no budget this quarter for a full monitoring pipeline. Ship the cheap version: a PM hand-reviews fifty flagged corrections a week in a spreadsheet, upgraded to real tooling once the pattern's proven.
The model got better, for real: say Frondwise's next version drops the real fern mix-up rate from 1 in 30 to 1 in 900. The three decisions don't change. A better model just moves where the needle sits on a rate you can now actually see, it doesn't excuse going back to trusting a single launch-day score.
Where people run it wrong.
They treat a patched ticket as proof the underlying rate is fixed, instead of proof that one case is fixed.
They build a denylist of banned outputs and call the risk closed, when the real failure was never naming the banned thing in the first place.
They write the acceptance bar as "must always be right," which sounds safer and measures nothing.
How to use it live. Ask one question before trusting any "it passed" for a probabilistic feature: "Passed compared to what, and who's still checking?" That question alone usually tells you whether you're looking at a real feedback loop or a score somebody measured once and framed.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Doesn't a hard-coded toxicity check make the feedback loop unnecessary?" Response: no, the toxicity check catches the one output that can't be undone, but it only works because someone keeps the toxicity table current. A feedback loop is what tells you a new regional plant needs adding to that table in the first place.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on What changes when the product is probabilistic
- #1 Name three product decisions that change when a feature's output is probabilistic rather than deterministic.
- #2 A traditional feature either works or has a bug. Explain why that framing breaks for an LLM feature.
- #3 What does 'correct' mean for a summarization feature? Give a definition your engineering team could test against.
- #4 QA files a bug that reads: the model gave a wrong answer once. How do you triage it?
- #5 Explain the difference between a defect and an acceptable error rate to a non-technical executive.
- #6 Why can you not write an acceptance criterion like 'the output must be accurate' for a generative feature?