CaseIntermediateModel Fluency & the AI PM Role / What changes when the product is probabilistic / #7

Describe how you would set a quality bar for a feature whose output is free text.

ORDER · an AI travel itinerary planner ranking what a free text quality bar checks first

Trailscript is Thornwyck Travel's engine for writing a full day by day trip plan in plain text: restaurants, hours, how to get between stops, all of it. Griselda Vellacott owns its quality bar. Five months in, the app's own dashboard still said 4.1 out of 5. Then Hepzibah Niemczyk, three weeks into the support desk, read a complaint ticket out loud and asked a question nobody on the team could answer.

The direct answer
Score free text output on three separate dimensions, factual accuracy, completeness, then style, and never blend them into one number. Gate what ships on factual accuracy alone, checked against a live source or a real timetable, because a wrong fact breaks trust in a way a clunky sentence never does. Build that gate from a small hand checked sample first, then widen it, and let style stay an ungated score that only affects polish.
Do this, in order
  1. Score accuracy, completeness, and style as three separate numbers, never one blended score.Why: a blended score lets a stack of easy style points hide two failed facts.
  2. Gate what ships on factual accuracy alone, verified against a live source, not a feeling.Why: a hallucinated closed restaurant is the one mistake a traveler can't undo trust for once they've walked there.
  3. Build the eval set from a small hand checked sample first, about 180 itineraries, not a brand new pipeline.Why: the raw material, complaint tickets and a modest sample, already exists and is cheap to check now.
  4. Flag any claim the model can't confirm recently as worth checking, instead of stating it with full confidence.Why: this is what stops thin coverage in newer cities from reading exactly as sure as the well covered ones.
  5. Check completeness against exactly what the traveler asked for, day by day.Why: a skipped request is recoverable with a quick fix, but only if someone is actually checking for it.
  6. Leave style as a soft, ungated score that never blocks a ship.Why: a clunky sentence costs almost nothing to fix, so gating on it spends the same care a real fact check needs elsewhere.

How to answer this, stage by stage

Nobody is grading whether you can name three things a quality bar might check. They're grading whether you know which one a traveler actually can't forgive.

1
Pin it to one app, one traveler facing decision, one person
Say it like this
"Let's ground this in one product. Trailscript is Thornwyck Travel's engine for writing a full day by day trip itinerary, restaurants, hours, how to get between stops, all in free text. Griselda Vellacott owns its quality bar. Amaranta Ashenfeld runs the support team that hears about it the moment a line in that text is wrong."
Why this works
An abstract "how do you handle quality" answer turns into a slogan fast. One real product keeps every claim something you'd actually have to defend.
2
Say the method out loud before naming a single check
Say it like this
"I'll run this as ORDER. Name what the bar is actually protecting, find which kind of mistake is hardest to walk back, say what has to exist before any number means anything, say what's cheap to check first, then give the actual order I'd tackle the three dimensions in."
Why this works
Two seconds of structure beats five checks arriving in whatever order they occur to you.
3
Name what the bar is actually protecting
Say it like this
"Here's what the bar is really for. It's not a quality score for its own sake. It's protecting the moment a traveler reads Trailscript's plan and just goes, instead of re-checking every line herself. The second she starts double checking hours and addresses on her own phone, the product has already lost, whatever the average star rating says."
Why this works
Without naming what's actually protected, ranking three dimensions is a preference dressed up as a method.
4
Find the dimension that can't be undone
Say it like this
"Three kinds of mistake can show up in that text: a clunky sentence, a skipped request, and a fact that's just wrong, a restaurant that closed, a train that doesn't run. Only one of those survives being discovered. A traveler forgives a stiff sentence in about the time it takes to read the next one. Someone who walks to a shuttered restaurant because Trailscript said it was open doesn't wait for an apology update. Forty four percent of them just don't book their next trip with us."
Why this works
This is the hardest step, and the one a rushed answer skips straight past to "check everything equally."
5
Say what has to exist first, and what's cheap to check
Say it like this
"None of these three scores mean anything yet, because right now Trailscript gets one blended one to five rating from a single reviewer. Before I write down a single number, I need two things: a rubric that scores accuracy, completeness, and style separately, and a small hand checked sample, about a hundred and eighty recent itineraries, checked against a real places source and real timetables, not against how the sentence reads. Both are cheap. The complaint tickets already exist, and a hundred and eighty is small enough to check by hand in a week."
Why this works
Naming the dependency is what stops an unproven bar from ever getting written down as policy.
6
Give the rank, defend the top pick, and close
Say it like this
"So, in order. Factual accuracy ships first, as a hard gate, and any claim the model can't confirm recently gets flagged instead of stated with full confidence. Completeness is checked second, against exactly what the traveler asked for. Style comes last, a soft score that never blocks a ship, because a clunky sentence costs a traveler ten seconds and a rewrite fixes it completely. Rank by what a traveler can't forgive, not by what's easiest to measure."
Why this works
Restates the direct answer with the actual order and the reason, so the interviewer leaves with the decision, not the story.

Let's learn

Trailscript writes a full trip plan in plain text: a named restaurant for lunch, a train time for the day trip, a walking route between the two, five or six days at a stretch.

Hand sketched labeled parts diagram titled What is actually inside one itinerary. A central document icon labeled One itinerary, with five labeled callouts around it: Named restaurant, Opening hours, Train connection, Walking route, and Tone and phrasing.
One itinerary reads like one piece of writing. It's actually five different kinds of claim stacked together, and only some of them are facts you can check.

Before Trailscript, a traveler put together a five day trip herself: a dozen browser tabs, a guidebook, a phone call or two to check a restaurant was still open. It took most of an evening. Trailscript writes the same plan in about ninety seconds, close to nine hundred of them a day now, about a hundred and thirty thousand across its first five months live.

Knowledge spark: what's a live source? A place Trailscript can check right now, a restaurant's own booking page, a transit agency's real timetable, instead of a guess dressed up as a fact. If a claim has no live source behind it, it isn't really known. It's just likely.

Complaint tickets tagged "itinerary was wrong" came in at about the same rate the whole time, 412 of them over those five months. Split by hand, 239 were factual, a closed venue, wrong hours, a train that doesn't run. 111 were completeness gaps, a skipped day, a stated request the plan ignored. 62 were style complaints, clunky phrasing, a cold or repetitive tone.

Follow up satisfaction score by complaint type, out of 5
5 2.5 0 no complaint logged, 4.6 1.8 Factual accuracy 3.1 Completeness 3.9 Style only
Factual accuracyCompletenessStyle only
Style complaints still land net positive, 3.9 out of 5. Factual complaints don't come close, 1.8, less than half the baseline for an itinerary with no complaint at all.

Here's the turn. The overall number never moved much, so nobody went looking. What moved was which kind of complaint made up the total, and that shift was invisible from the dashboard alone.

The dashboard said 4.1 out of 5 the whole time. It never said which promise was breaking.
Hand sketched left to right flow diagram titled How the gap got wide enough to notice. Four connected boxes reading: 40 cities, 300 more join, Data runs thin, this box emphasized in red, Text sounds sure.
Trailscript launched validated in 40 cities with data someone had actually checked. Three hundred more joined on a faster, rougher feed, and the writing never got any less confident about it.
Share of complaint tickets that were factual accuracy errors, week by week
70% 35% 0% week 8, coverage jumps to 340 cities wk2 wk6 wk10 wk14 wk18 wk20, Hepzibah asks, 58%
Weekly share, factual accuracy tickets
35 percent of complaints in week 2. 58 percent by week 20. The blended star rating never showed this climbing, because the total ticket count barely moved, only its makeup did.

What it costs at its worst: a traveler stops trusting the plan and starts re-checking every line herself, which is exactly the evening of work Trailscript was supposed to save her. Multiply that by three hundred cities where the underlying data was never as solid as the forty it launched on, and the product quietly gets worse in exactly the cities it was built to help most.

The decision that mattered Trailscript's dashboard gave every itinerary one blended score, one to five stars, from a single reviewer. That was fine when the team reviewed itineraries by feel and caught problems anyway. It stopped being fine the day coverage grew faster than the data behind it, and nobody had a way to see the gap opening until a new hire happened to read the right ticket.

What I would leave alone: Trailscript's short one day city breaks in the original forty flagship cities. Coverage there is dense and checked often, so the accuracy risk barely exists. Building the same confidence gate machinery for those itineraries would just slow down cases that were never actually broken.

The lesson: a quality bar that gives one number can't tell you which promise it's actually keeping. Split the score before you trust it, or you find out which promise broke only after a traveler already has.

Now here is the same thing as a story

Read the short version above when you're in the room. Read this one when you want to feel why a closed restaurant, not a clunky sentence, was the thing that could have cost Thornwyck the trip after next.

Griselda Vellacott built Trailscript's first version in forty cities, the ones Thornwyck already had years of curated data on: real booking pages, real transit timetables, restaurant hours somebody had actually called and confirmed.

For the first three months, that carefulness showed. A traveler typed in five days in Lisbon and got back a plan with a named lunch spot for day two, a train time for the day trip to Sintra, a walking route between the two. Complaints trickled in, the odd typo, a museum closed on a Monday nobody had flagged. Nothing that shook anyone's confidence in the product.

Then Thornwyck asked Trailscript to cover everywhere Thornwyck sold trips, not just the forty flagship cities. Three hundred more went live over six weeks. The new cities didn't come with the same kind of data behind them. Trailscript pulled from a faster, rougher feed instead, one that kept listing a restaurant as open for months after it had actually closed.

Nobody decided this was risky on purpose. On paper, the rollout looked like more of the same good thing.

The complaint tickets kept arriving at roughly the same rate they always had, so the dashboard's blended score barely moved, 4.1 out of 5, week after week. What the dashboard couldn't show was what kind of complaint was landing. Factual complaints, a place that was shut, a train that didn't run, made up about a third of tickets in week two. By week twenty, they were well over half.

Hepzibah Niemczyk had been on the support desk three weeks. Reading through a batch of recent tickets during onboarding, she found one where a traveler had walked twenty minutes to a restaurant Trailscript called the best table in the neighborhood, only to find it closed, permanently, for two years. She brought it to Griselda with a plain question. "How do we know the places Trailscript names are actually still open?"

Hand sketched comparison diagram titled What Hepzibah found versus what the dashboard showed. Left panel, a person icon labeled Hepzibah's question, caption How do we know the places are still open. Right panel, a gauge icon labeled The dashboard, caption 4.1 out of 5, looks fine.
Two things that were both true at the same moment, and only one of them was useful to Griselda.

Nobody on the team had a real answer. Not because nobody cared. Because the only score Trailscript had ever kept, one to five stars, blended by a single reviewer, couldn't separate a wrong fact from a clunky sentence. Both just pulled the same number down a little.

Griselda pulled ninety days of tickets and split them by hand into three piles: wrong sentence, missing request, wrong fact. The wrong-fact pile hurt the most, and not by a little. Itineraries with a factual complaint scored 1.8 out of 5 on the follow-up survey. Style-only complaints still scored 3.9, mildly annoying, forgotten by the next trip. Forty four percent of travelers who hit a wrong fact didn't book their next trip through Thornwyck within ninety days. Only seven percent of the style-only group did the same.

That gap was the real finding, not the closed restaurant by itself. A clunky itinerary is a rough first draft. A wrong fact is a promise Trailscript can't take back once someone's already walked there.

The decision Griselda would take back sat right at the start, before Trailscript covered a single new city. The team built one dashboard score because it was faster to read and simpler to build than three. That was a fine call for forty cities Thornwyck already knew cold. It stopped being fine the moment coverage grew faster than the data backing it, and nobody had a way to see the gap opening until a new hire happened to read the right ticket.

Run the rollout again, with three scores from day one instead of one. The same three hundred new cities go live over the same six weeks. This time, the accuracy score drops on its own, fast, days after the rougher feed starts feeding in, not twenty weeks later when a new hire stumbles onto it by accident. Griselda pulls the accuracy number in week nine, not week twenty, and holds the newest cities to category-level suggestions until the feed catches up.

One design let a wrong fact hide behind a decent-looking average for five months. The other would have caught it in nine weeks.

What I'd tell myself, back when we shipped one blended score because it was simpler: simple was the right call for the size of team we were then. It just wasn't a call built to survive the day we started covering three hundred more cities than anyone had actually checked.

ORDER: what a free text bar has to protect first

Not a way to make three numbers sound rigorous. ORDER is what forces you to say, before any bar ships, which broken promise a traveler actually can't forgive.

OOutcome. What is the bar actually protecting?
Not "make Trailscript's writing better." One specific thing: the moment a traveler reads the plan and just goes, instead of re-checking every line herself. Every candidate check, tone, grammar, missing days, wrong facts, gets judged against whether it protects that moment, not against how easy it is to measure.
Without naming what's actually protected, ranking accuracy above style is a preference, not an argument.
RReversibility. Which kind of mistake can't be walked back?
A wrong fact. Forty four percent of travelers who hit a hallucinated closed restaurant or a phantom train connection don't book their next trip within ninety days, against seven percent for a style-only complaint and eighteen percent for a missed request. A clunky sentence gets forgiven the moment it's rewritten. A wrong fact already cost someone twenty minutes and a walk before anyone can rewrite anything.
This is the hardest step, and the one a rushed answer skips straight past to "check everything equally."
Hand sketched comparison diagram titled Which mistake can be walked back. Left panel, a document icon in green labeled Clunky sentence, caption a rewrite fixes it completely, trust resets. Right panel, a question mark box icon in red labeled Wrong fact, caption trust does not reopen once someone has walked there.
One door swings back open the moment someone edits the sentence. The other stays shut, because the traveler already made the trip.
DDependency. What has to exist before any number means anything?
Two things, neither built yet: a rubric that scores accuracy, completeness, and style separately instead of blending them, and a labeled eval set, itineraries checked against a real source, a live places listing or a real timetable, not against how the sentence reads.
Naming the dependency is what stops a bar from getting written down before it's actually been checked against anything real.
Hand sketched flow diagram titled What has to exist before any bar means anything. Five boxes connected by wobbly arrows: Split into three, Hand check 180, Confirm the rate, this box outlined in blue to mark the step everything else depends on, 60 day window, Ship the gate.
The order a gut call skips: check the model against real cases before any bar built on top of it gets written down.
EEvidence. What's cheap to check before building the whole pipeline?
A hand checked sample of about a hundred and eighty recent itineraries, checked against a live places source and real timetables, plus the ninety days of complaint tickets that already exist. Neither needs new infrastructure. Both are ready in about a week.
Cheap evidence beats a bar that's only ever been tested on how confident it sounds.
RRank. State the order, defend the top pick.
Factual accuracy ships first, as a hard gate: any claim Trailscript can't confirm against a live source within sixty days gets flagged as worth checking instead of stated outright, or drops to a category level suggestion. Completeness ships second, checked against exactly what the traveler asked for. Style ships last, a soft score that never blocks a launch, because rewriting a clunky sentence costs almost nothing next to what a wrong fact costs.
If the rank would look the same with a different outcome named in step one, it was ranked by gut and the outcome got written afterward.

Three things worth stating directly, since this is where the real judgment sits. The alternative Griselda's team piloted, and rejected, was one master checklist, twenty items, accuracy and completeness and style all folded into a single pass by one reviewer. Fifty itineraries into the pilot, some were passing with seventeen of twenty checks clear while failing two of three fact checks outright, because a blended checklist can average away exactly the failure that matters most. It lost, for the same reason the old one to five score lost. The AI-specific failure worth naming by name is a grounding gap dressed as good news: the faster feed behind the new three hundred cities lists a closed restaurant as open for an average of fourteen months after it actually shuts, and a model that states a claim with full confidence has no way to say it's less sure about this one unless something forces it to check its own source's age first. The guardrail is that sixty day confirmation window, any named claim resting on a source older than that, or with no confirming source at all, gets flagged instead of stated flat. And the trade-off is real and accepted on purpose. Itineraries in newer, thinner cities read a little less specific, a well reviewed noodle place near the station instead of a name and an address, and each verified claim adds a small amount of generation time. That's traded away on purpose, because the alternative is finding out about the next closed restaurant from a support ticket instead of a flag that cost a few extra seconds and nothing else.

Hand sketched labeled parts diagram titled Three needles, not one. A central gauge icon labeled Trailscript's bar, with three labeled callouts around it: Accuracy ships first, Completeness ships second, Style ships last.
The whole answer to this question, in one picture. Three needles that move on their own beat a single needle that can hide any of them.

And if you want to be sure it really works, try it somewhere else

Same five letters, a veterinary chain instead of a travel app, and this time the fragile claim isn't a restaurant's hours. It's a dosage line.

Mendscript drafts the after visit instructions Elderfen Veterinary Group's fourteen clinics hand pet owners on the way out the door: medication name, dose, how often, what to watch for, when to call back. Ferdinandine Sorbier owns its rubric. For seven months, one staff reviewer rated each note one to five, mostly on whether it read clearly and matched the visit type.

The decision Elderfen's team would take back Mendscript's rubric scored a discharge note the same way Trailscript once scored an itinerary, one blended number from a single reviewer. That made sense when the clinics were small enough for a reviewer to catch a wrong dose on instinct. It stopped being fine the day the group scaled past what one reviewer's instinct could reliably cover.

Dr. Evangelina Loxbury caught the near miss during a routine chart review: a printout for a forty two pound dog recovering from a neuter had pulled its pain medication dose from a cached template built for a much smaller weight bracket, understating the dose by half. Nothing had gone wrong yet. The owner hadn't left the building.

Same rank, different lever, mapped straight onto ORDER: the outcome here isn't "a well written note," it's an owner who follows the instructions exactly as written because they trust them. Reversibility splits the same way: a wrong dosage or a missing warning sign can't be undone once a pet's already been given the wrong dose at home, while a stiff, clinical sentence gets fixed by any follow-up phone call. Dependency is the same shape too: verify every dosage and warning line against the actual prescription record before scoring anything, not against how professional the note sounds. Evidence is a hand checked sample of the last sixty discharge notes against real prescription records, cheap, and already sitting in the system. And the rank lands the same way: dosage and warning claims verified against the prescription record first, as a hard gate, every required warning category present second, bedside tone last, worth polishing, never worth blocking a discharge over.

Hand sketched decision tree titled Mendscript's discharge note, ranked the same way. Root box reads One discharge note, three kinds of claim, branching into three outcomes: dosage and warnings vs the prescription record leading to Hard gate, ships first, every required warning category present leading to Checked second, and bedside tone and readability leading to Polished last, never blocks.
Different animal, same order. The claim a pet owner can't safely act on wrong ships first, not the sentence that only needs to sound kind.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the rank: verify dosage against the prescription record first, check required warnings second, polish tone last.
Cost: there's no budget to check every note this quarter. Check the highest risk categories first, pain medication and anything with a dosage number, before the rest.
The model got better, for real: say Mendscript's dosage accuracy climbs from occasional misses to near perfect. The rank doesn't move. Dosage verification still leads, it just clears the gate faster.

Where people run it wrong.
They score everything on one blended scale and never find out which promise is actually breaking.
They fix the loudest complaint instead of the one nobody can undo.
They wait for a serious incident to happen before splitting the score, instead of splitting it before anything happens.

How to use it live. Ask, before ranking anything: "which of these mistakes can the person fix themselves after the fact, and which one already happened by the time anyone notices?" That question sorts the three dimensions faster than any framework does on its own.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework is this, and what's its one job?
Tap to flip
ANSWER
ORDER: rank the dimensions of a quality bar by which one is hardest to walk back if it's wrong. Built for prioritization questions, including what a free text quality bar checks first.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Griselda Vellacott, who owns Trailscript's quality bar at Thornwyck Travel, an AI tool that writes full day by day trip itineraries in free text.
3 · THE OUTCOME
What is the quality bar actually protecting?
Tap to flip
ANSWER
The moment a traveler reads the itinerary and just goes, instead of re-checking every line herself. Not a quality score for its own sake.
4 · THE DIMENSION THAT CAN'T BE UNDONE
Which kind of mistake is hardest to walk back, and what's the proof?
Tap to flip
ANSWER
A wrong fact, a restaurant Trailscript calls open that closed two years ago. 44 percent of travelers who hit one don't book their next trip through Thornwyck, against 7 percent for a style-only complaint.
5 · THE OLD DECISION
What decision would Griselda take back?
Tap to flip
ANSWER
Scoring every itinerary with one blended one to five rating from a single reviewer, instead of separate scores for accuracy, completeness, and style. It made sense for forty curated cities reviewed by feel; it hid exactly which dimension was breaking trust once coverage tripled.
6 · THE NUMBER
Fill in the blank: factual accuracy complaints scored ___ out of 5 on the follow-up survey. Style-only complaints scored ___ out of 5.
Tap to flip
ANSWER
1.8 out of 5 for factual accuracy, 3.9 out of 5 for style-only. Style complaints stayed net positive; factual ones didn't come close.
7 · THE RANK
Same three dimensions, what's the actual order, and what changes?
Tap to flip
ANSWER
Factual accuracy first, a hard gate against a live source. Completeness second, checked against what the traveler asked for. Style last, a soft score that never blocks a ship.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which one, and what's the equivalent fragile claim?
Tap to flip
ANSWER
Mendscript, Elderfen Veterinary Group's AI-drafted pet discharge instructions. The fragile claim there is a wrong medication dosage or a missing warning sign, not a stiff bedside tone.

Check yourself Score: 0 / 0

Multiple choice
1. Which complaint type would do the most lasting damage to Thornwyck if it went unfixed?
  • A. Repetitive sentence openers across a few days of the plan.
  • B. A restaurant Trailscript calls open that actually closed two years ago.
  • C. A walking route described in slightly stiff, clinical wording.
  • D. An itinerary that runs about half a page longer than usual.
Show hint
Check the reversibility numbers, the non-rebooking rate by complaint type.
Show answer
B. A wrong fact drove a 44 percent non-rebooking rate, against 7 percent for a style-only complaint. The other three options are style or format issues, cheap to fix and quickly forgiven.
Fill in the blank
2. Trailscript launched validated in ___ cities. Coverage later expanded by 300 more to reach ___ cities total, using a faster feed for venue status in the newer ones.
Show hint
Check the flow diagram in "Let's learn" and the story's opening paragraphs.
Show answer
40 cities, then 340 total. That gap in data quality, not a sudden model failure, is the real reason factual complaints climbed.
True or false
3. True or false: Trailscript's blended dashboard score of 4.1 out of 5 stayed roughly flat for months even while the share of complaint tickets that were factual accuracy errors climbed from about 35 percent to 58 percent.
  • True
  • False
Show hint
Look at the line chart in "Let's learn" and the paragraph right before it.
Show answer
True. The total number of complaints barely moved, so the blended average barely moved. What changed was invisible inside that one number: which kind of complaint was landing.
Short answer, name the old decision
4. What old decision would Griselda take back, and why did it make sense when it was made?
Show hint
Look at the key point box titled "The decision that mattered" in "Let's learn."
Show answer
Model answer: Giving every itinerary one blended one to five score from a single reviewer, instead of scoring accuracy, completeness, and style separately. It made sense for the original forty curated cities, reviewed by feel, because the team could catch problems anyway. It stopped making sense once coverage tripled and nobody had a way to see which dimension was actually breaking.
Short answer, apply it yourself
5. Think of a free text AI output you've used yourself, a summary, a draft email, a set of directions. What's one factual claim in it you'd want checked before you trusted it, and one stylistic thing you wouldn't bother gating on?
Show hint
Split what the output claims into "can be checked against a real record" and "just a matter of taste."
Show answer
Model answer: An AI meeting summary tool. Check whether a stated action item's deadline actually appears in the transcript, not invented by the model. Don't gate on whether the summary opens each bullet with a different verb.
Short answer, work the number
6. If Thornwyck's factual accuracy complaint share had stayed at week two's 35 percent instead of climbing to 58 percent by week twenty, roughly how many fewer of the 412 total tickets would have been factual accuracy complaints?
Show hint
Find 58 percent of 412 and 35 percent of 412, then subtract.
Show answer
About 95 fewer tickets. 58 percent of 412 is about 239. 35 percent of 412 is about 144. The gap, roughly 95 tickets, is what three hundred cities of thinner data actually cost.
Before you close the answer
Why this works
Tests whether you'll treat "set a quality bar" as a real ranking problem with a named winner and loser, not a checklist, and whether you know a wrong fact and a clunky sentence are genuinely different kinds of failure with different real costs.
Follow-up traps
"Why not just have a person review every single itinerary before it ships?" Response: at close to 900 itineraries a day, full human review isn't sustainable, and it still wouldn't tell you which dimension was failing, only that something felt a little off, which is the exact problem the blended score already had.

"Isn't flagging low confidence claims going to make every itinerary in a new city look worse?" Response: yes, on purpose. An itinerary that admits it isn't sure about one detail is more trustworthy than one that's confidently wrong, and that trade is worth making every time.
If pressed
The sixty day confirmation window wasn't picked in the abstract. Griselda's team found the faster feed behind the newer three hundred cities lists a closed restaurant as open for an average of fourteen months after it actually shuts, which is the real mechanism behind the hallucinated-restaurant problem, not a one-off model slip. Sixty days was chosen as the point where a stale listing is still more likely right than wrong, without waiting for a fix that could take over a year to show up in the feed on its own.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more