ConceptAdvancedQuality, Cost & Token Economics / Leading vs lagging indicators for AI / #13

Describe a leading indicator specific to agent products.

Find the number that moves before Wend's real booking success rate does, and don't let a quiet week convince you the problem is fixed.

The direct answer
Watch the rate Wend pauses to ask a person before it finishes a booking it used to handle alone. A steady rise in that confirmation request rate shows up two to six weeks before the real booking success rate ever drops, while the completion number on the dashboard still looks fine. Never read a falling confirmation rate as good news by itself. Pair it with a same day audit of what actually happened on the trip, because an agent tuned to stop asking looks more autonomous on that one tile while it quietly starts booking the wrong thing.
Do this, in order
  1. Put Wend's confirmation request rate on the weekly review, next to completion rate, not behind it.Why: it moved from three percent to fourteen percent over six weeks while completion rate held near ninety one percent the whole time.
  2. Never let a falling confirmation rate count as a fix on its own.Why: the rate and the agent's real accuracy can move in opposite directions the moment someone touches the same setting that produces both.
  3. Run a same day audit against a fixed set of past bookings whenever the confirmation rate moves more than five points in a week, in either direction.Why: this is the one check that tells a real improvement apart from a quieted alarm.
  4. Slice a rising confirmation rate by task type and partner before changing any setting.Why: at Journeo the whole six week rise traced back to one new hotel partner's data, not the entire agent losing its footing.
  5. Log the agent's mid task replanning rate as a second signal, not the main one.Why: replanning also rises before a real slip, but it's noisier for an agent that's supposed to replan when a fare changes, so it backs up the confirmation rate instead of replacing it.
  6. Give one named person the power to freeze the confidence threshold the same day a floor gets crossed.Why: a number nobody can act on the same day is just a chart on a wall.

How to answer this, stage by stage

Nobody is grading whether you can name a metric. They are grading whether you can catch a healthy looking tile lying to you before a customer does. Seven moves get you there.

1
Scope it to one real agent, one real owner
Say it like this
"Let's ground this. Journeo runs Wend, a chat agent that searches flights and hotels and books both on its own, no clicking required. Demola Ferran leads the product team that owns Wend's weekly numbers."
Why this works
Gives the interviewer a real product before you name a single metric.
2
Say your structure out loud
Say it like this
"I'd run this through LEAD. Link it to the outcome that actually matters, find the early signal, say plainly how that signal gets gamed, then say what I'd do at each threshold."
Why this works
Two seconds of structure, so the interviewer knows where you're headed before you get there.
3
Name the outcome, and say why the obvious number can't be trusted alone
Say it like this
"The number that matters is whether a trip finishes right with nobody stepping in to fix it after. Not whether Wend says a booking went through. Wend can mark itself done the second it books, before anyone's actually checked in anywhere."
Why this works
Separates the model's own self-report from the outcome a person actually lives through, which is the whole reason a real audit has to exist.
4
Give the early signal, the actual answer to the question
Say it like this
"The signal that moves first is the rate Wend stops and asks a person before finishing something it used to just do, like picking a substitute room when the one you wanted is sold out. At Journeo that rate climbed from three percent to fourteen percent over six weeks, while completion rate hadn't budged."
Why this works
This is the direct answer, stated with a real number, not a definition.
5
Say plainly how the signal gets gamed
Say it like this
"Here's the trap. If someone just turns down how easily Wend asks, the confirmation rate drops right back down, and the tile goes green again. But nothing about the agent's real judgment improved. It's just staying quiet about the same calls it used to check."
Why this works
Naming the abuse case unprompted tells an interviewer you've actually lived with this metric, not just defined it.
6
Give the decision at each threshold, in real numbers
Say it like this
"Inside its normal band, two to five percent, with the weekly audit under one percent wrong, I leave it alone. Past six percent for two straight weeks, I slice it by task type before I touch anything. Any sudden drop right after a threshold change gets treated as unproven until that week's audit clears it."
Why this works
Turns the metric from a chart into something a person actually does something about.
7
Close on what you ruled out, and what the fix costs
Say it like this
"We looked at asking Wend to confirm every single booking, and ruled it out. That kills the whole point of an agent that books in ninety seconds, because now a person has to be awake and answering. The real cost of the fix we kept is that the audit itself now runs daily instead of monthly, and any threshold change has to sit for a day before it's allowed to count as a win."
Why this works
A rejected option and a named cost is what makes this a decision instead of a wish that all three were free.
If you remember one thing A metric that goes quiet can mean the problem got fixed, or it can mean somebody stopped measuring it out loud. From the tile by itself, you can't tell which. By the time you can, the trip already happened.

Let's learn

On a TV bolted to the wall in Journeo's ops bay, one tile updates every morning: the share of Wend's bookings that finished without anyone stepping in to fix them.

Wend is a chat agent inside Journeo's app. Tell it your dates, your budget, and where you're headed, and it searches flights and hotels, picks the best match, and books both, with nobody clicking a single button.

Before Wend could book on its own, a person still picked from a results list and confirmed every leg by hand. That took about eighteen minutes for a return flight with a hotel attached, most of it spent comparing fares that all looked about the same. Now Wend finishes the whole thing itself in under ninety seconds, for roughly nine bookings out of every ten it touches. The tile on the wall has read close to ninety one percent for months.

Wend's confirmation request rate, by week
16% 0% threshold tuned down 14% 3% 4% wk1 wk3 wk5 wk7 wk9
The rate crept up for six weeks straight before anyone acted on it. In week seven someone finally acted on it, just not the way you'd want.

Here is the turn. That ninety one percent never moved. Not once, for six straight weeks. And that steadiness is exactly the problem, because underneath it, the rate Wend paused to ask a person before finishing a booking climbed from three percent to fourteen percent, and the completion number was about to fall for real, not because Wend got worse on the tile, but because the tile was never built to notice.

The tile didn't lie. It just never had a way to tell the difference between fixed and quiet.
Knowledge spark: what is a self-reported success flag? The tag an AI agent puts on its own work the moment it finishes acting. Booked, done, no problem. It comes from the agent watching itself, not from anyone checking what actually happened on the trip. It can be honest for months and still be the wrong thing to trust the week the agent's own judgment shifts.

At its worst, this doesn't look like an outage. It looks like nothing. A traveler checks into the wrong room type, or gets an extra night nobody asked for, sorts it out at the front desk, and never opens a support chat, because as far as Wend's own record shows, nothing went wrong.

Booking completion, weeks seven to nine: what the tile said vs. what the audit found
100% 0% 90% 78% Live tile (self-reported) Post-trip audit (real outcome)
Wend's own tile still read close to ninety percent, because it marks a booking done the moment it books, not after the trip happens. The post-trip audit, run on a delay, found the real number.
The decision that mattered Letting the same threshold that decides how often Wend asks also decide, by itself, whether the confirmation-rate alarm stays quiet. That was fine while nobody had a reason to touch it. It stopped being fine the week someone did.

What I would leave alone: Wend's confirmation rate for small stuff, window seat or aisle, doesn't need a floor at all. Getting that wrong costs someone one annoyed email, not a wrecked trip. Only the asks that stand between the agent and money actually leaving someone's card are worth watching this closely.

The lesson: a number that goes quiet can mean the problem got fixed. It can also mean somebody stopped measuring it out loud. Those look exactly the same on a wall-mounted tile, and the only way to tell them apart is to check something the tile itself doesn't control.

Now here is the same thing as a story

The short version sits above. Read on for the Tuesday a customer's thank you note was the only warning anyone got.

Demola Ferran has led product for Wend for two years, since before it could book anything on its own. Back then it only searched, and a person still clicked confirm on every leg by hand.

For most of a year after Wend started booking on its own, the tile in the ops bay barely moved. Ninety, ninety one, ninety two percent, back to ninety one. A boring number was the whole point. It meant Wend was doing today what it did yesterday.

The confirmation rate crept up so slowly nobody called a meeting about it. Three percent in week one. Four in week three. Six in week four. By week six it sat at fourteen, and the completion tile still read ninety one, steady as ever, because the trouble hadn't reached a real trip yet.

Then came a Friday afternoon in week seven. An engineer on call that week, watching the confirmation rate graph climb on a shared screen, did the fastest thing available to make the number stop climbing: turned the confidence threshold down a few points, so Wend would ask less often. The graph flattened out. He logged it as handled and went home for the weekend.

The completion tile never moved. It couldn't. It was reading what Wend said about itself, and Wend, asking fewer questions now, said done more often than not.

Two weeks later, Iona Corry opened a support chat from a hotel lobby in Lisbon, mid-trip with her family. Not to complain. To say thanks. Wend had booked her a room a size smaller than what she'd asked for, and the front desk had sorted it out at check-in with no trouble. "No big deal," she wrote. "Your app usually asks before it does anything like that though. This one it just went ahead."

Nobody flagged it. A five-star chat doesn't get read twice. It sat in a queue until a teammate, pulling transcripts for an unrelated project, noticed the line about Wend usually asking first and thought that was an odd thing for a happy customer to mention.

Demola pulled the real numbers that afternoon. The confirmation rate had been sitting near four percent for two weeks, which read as a win to anyone glancing at the tile. The post-trip audit, which only ran once a month, hadn't caught up to any of it yet. When it finally did, three weeks after the Friday the threshold got turned down, the real completion rate for that stretch came back at seventy eight percent.

We did not lose one hotel room in Lisbon. We lost the three weeks between a fake fix and finding out it was fake.

Nobody at Journeo could say how many other trips like Iona's had gone quietly wrong in that window, because none of them had generated a ticket. A happy guest at a front desk leaves no record built to catch this.

Here is the part that actually mattered. Demola's team never had a real number for what happened on a trip in real time. They had a flag Wend set about itself, and a flag can be told to say anything the thing measuring it wants it to say.

Setting up that alert eighteen months earlier had taken about ten minutes in a stand-up. One number, one alarm, one threshold. It was the right call. Nobody could have known that the fastest way to quiet a rising alarm would be to turn down the very knob that produced it, and hand that knob to whoever happened to be on call that Friday.

Run the same six weeks through the fixed design. The audit runs same day now, any time the confirmation rate moves more than five points in a week. Week seven's drop gets flagged within hours, not caught by a customer's thank you note three weeks later. The threshold change gets reverted before Iona ever books her trip, and the real completion rate never leaves ninety one.

One design let the alarm get turned down by the same hand that set it off. The other makes that exact move show up as its own event, checked before it's allowed to count as fixed.

What I'd tell myself, back in that ten-minute stand-up: the day you build an alarm, ask who has the power to make it quiet, and whether quiet and fixed will ever look different to them.

LEAD, and the tile that couldn't tell fixed from quiet

This is a metric question, so LEAD does the work: find the outcome, find what moves first, name how it gets gamed, then say what you'd actually do.

L
Link. The outcome that actually matters.
Not whether Wend's own flag says booked. Whether the trip that actually happened matched what the traveler asked for, checked by something other than the agent itself.
In this answer: booking completion rate from the post-trip audit, not the live self-reported tile.
E
Early signal. What moves before the outcome does.
Wend's confirmation request rate, the share of bookings where it stops and asks a person before finishing something it used to do alone.
Climbed from three percent to fourteen percent over six weeks while completion rate sat still the whole time.
A
Abuse. How the signal gets gamed.
Turn down the same confidence threshold that makes Wend ask, and the confirmation rate drops right back to a healthy-looking number, with nothing about the agent's real judgment improved.
This is exactly what happened at Journeo in week seven. The tile read fixed. The audit, three weeks later, read seventy eight percent.
Hand sketched two panel comparison titled A quiet dashboard is not the same as a fixed agent. Left panel, a gauge icon labeled the tile, caption confirm rate down, tile reads green. Right panel, a person icon labeled the trip, caption wrong hotel booked, nobody asked.
Same week, two very different truths. One of them is easy to check without asking the agent's permission.
D
Decision. What you'd actually do at each threshold.
A number nobody acts on is decoration. Every band on this metric maps to a specific move, not a vague "keep watching it."
Normal band, two to five percent, audit under one percent wrong: leave it alone. Past six percent for two weeks: slice by task type before touching anything. A sudden drop right after any threshold change: unproven until that week's audit clears it.

Three things worth naming directly, since this is where the real judgment sits. We considered leading with the agent's own mid-task replanning rate instead of the confirmation rate, and set it aside as the lead metric: Wend is supposed to replan when a fare changes mid-search, so a rising replan rate is often healthy, and it's a harder number for anyone outside engineering to look at on a Tuesday and act on. We kept it running as a second, corroborating check instead. The failure worth naming by name is a form of confident wrongness: the model's own confidence calibration shifted after one hotel partner changed the shape of its availability data, so Wend started treating genuinely uncertain room-type matches as safe to book alone. The guardrail is a task-type-sliced golden set, a batch of past bookings with known right answers, checked against the confirmation rate every week rather than trusting the tile by itself. None of this is free. Asking more often is slower, and it can lose the booking entirely if nobody answers in time, which trades against the whole reason Wend exists, a flight and a hotel booked in under ninety seconds with nobody clicking anything. The bar we set is a calibrated one, not a hard rule: the confirmation rate and the golden-set wrong-action rate should move together, most weeks, within their normal band. When they stop moving together is the signal, not any single number on its own.

And if you want to be sure it really works, try it somewhere else

Same four letters, a warehouse reorder agent instead of a travel agent, and the trap is the same one dressed differently.

Northbank Supply runs Tally, an agent that watches stock levels across its warehouses and places reorders with vendors on its own, inside a spending cap, no purchase order needing a person's signature. Nkem Segovia runs procurement operations there.

L, link. Not whether Tally's own log says order placed. Whether the order that actually arrived matched what was needed, right vendor, right quantity, right price, checked against the receiving dock's own count, not Tally's own record.
E, early signal. The rate Tally pauses to ask a person to confirm an order before it submits one, on purchases it used to place alone under its spending cap. It rose from two percent in week one to eleven percent by week four, while the dock's error count stayed low the whole time.
A, abuse. In week five, procurement raised Tally's auto-approve spending cap so it would stop asking as often. The confirmation rate fell to two percent by week six. The dock's own count of wrong orders kept rising the entire time, from one wrong order a week to six.
D, decision. Same shape of table. Inside the normal band with the weekly dock audit clean, leave it. Rising for two weeks straight, slice by vendor before touching the cap. Any fast drop right after someone raises the cap: unproven until that week's dock count clears it.

Tally's confirmation rate vs. what the dock actually found, weeks three to six
12 0 wk 3 wk 4 wk 5 wk 6
Confirmation rateDock error count
The confirmation tile looks like a steady improvement after week four. The receiving dock's own count, checked separately, moved the other way.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip to the fix: whatever leading signal your dashboard shows, pair every reading of it with an independent check that doesn't share a setting with the signal itself.
Cost: engineering says a same-day audit is a month out. Don't read the live tile alone as a stopgap in the meantime, pull a hand sample of last week's bookings once a day until the real pipeline ships.
The model got better, for real: say Wend's underlying model genuinely improves next quarter. That still isn't the same claim as the confirmation rate falling for a good reason. A better model can lower that rate honestly, which is exactly why the number needs a second check that doesn't just take its word for it.

Where people run it wrong.
They watch the leading number and the lagging number on separate dashboards, so nobody ever puts them side by side and asks why one moved without the other.
They let whoever can change the underlying setting also be the one who reads the alert it produces.
They wait for a monthly audit to confirm what a weekly one would have caught in days.

How to use it live. Say the reframe before naming a single metric: "a number that used to predict trouble can be tuned to stop predicting it, without the thing it was watching ever getting better." That buys you room to give the real answer instead of reciting "track leading indicators" on reflex.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
Which framework fits a metric question like this one?
Tap to flip
ANSWER
LEAD: link to the outcome that matters, find the early signal, name how it gets gamed, decide what to do at each threshold.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Demola Ferran, who leads product for Wend at Journeo, and Iona Corry, the traveler whose thank you note was the only warning anyone got.
3 · WHAT STOPPED
What did Demola's team stop doing once the tile looked steady?
Tap to flip
ANSWER
They let a monthly audit be the only check on what actually happened on a trip, and let Wend's own self-reported flag stand in for the real outcome in between audits.
4 · THE SIGNAL
What early signal moved before completion rate did, and by how much?
Tap to flip
ANSWER
Wend's confirmation request rate, the share of bookings where it stopped to ask a person first. It rose from three percent to fourteen percent over six weeks.
5 · THE OLD DECISION
What decision would you take back, and why did it make sense at the time?
Tap to flip
ANSWER
Letting the same confidence threshold that sets how often Wend asks also be the only thing that could quiet its own rising-rate alarm. It was fine while nobody had a reason to touch it.
6 · THE NUMBER
Fill in the blank: over six weeks, Wend's confirmation rate rose from ___ percent to ___ percent, while the completion tile stayed near ___ percent the whole time.
Tap to flip
ANSWER
3 percent; 14 percent; 91 percent. The rise was real. The tile never showed it.
7 · THE REPLAY
Same six weeks, new design, what changes?
Tap to flip
ANSWER
A same day audit fires the moment the confirmation rate moves more than five points in a week. The week seven threshold change gets caught and reverted within hours, not found three weeks later in a customer's thank you note.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which one, and what's the parallel?
Tap to flip
ANSWER
Tally, Northbank Supply's reorder agent. Same LEAD letters, a purchase-order confirmation rate instead of a booking one, gamed the same way by raising the auto-approve spending cap.

Check yourself Score: 0 / 0

Multiple choice
1. What is the real problem with reading Wend's completion tile as proof the agent is healthy, on its own?
  • A. It updates too slowly to be useful day to day.
  • B. It comes from Wend's own self-report, not from checking what actually happened on the trip.
  • C. It only counts flight bookings, not hotels.
  • D. It can't be shown on a wall-mounted screen.
Show hint
Think about who marks a booking done, Wend or a person.
Show answer
B. Wend flags itself done the moment it books. The real outcome only shows up once someone checks the trip against what actually happened.
True or false
2. True or false: Wend's confirmation request rate and its completion rate moved together over the six weeks in this story.
  • True
  • False
Show hint
Check what the tile on the wall read during those six weeks.
Show answer
False. The completion tile held steady near ninety one percent the entire time, while the confirmation rate quietly climbed from three to fourteen percent underneath it.
Fill in the blank
3. Over six weeks, Wend's confirmation request rate rose from ___ percent to ___ percent.
Show hint
Check the first chart in "Let's learn."
Show answer
3 percent to 14 percent. That rise, not the completion number, is what predicted the real trouble weeks before it hit.
Multiple choice
4. Why couldn't Demola's team just treat the confirmation rate falling back to four percent in week eight as good news?
  • A. Because a lower confirmation rate is always the wrong direction for this metric.
  • B. Because the drop came from turning down the same threshold that produces the metric, not from the agent getting more accurate.
  • C. Because Wend can't report numbers below five percent accurately.
  • D. Because the drop only applied to hotel bookings, not flights.
Show hint
Think about what actually changed in week seven.
Show answer
B. The rate fell because someone lowered the threshold that decides how often Wend asks, not because Wend got better at knowing when to ask.
Short answer, apply it yourself
5. Pick a product you use that acts on your behalf without asking every time. What's one leading signal that product could watch, that would move before the thing you'd actually notice going wrong?
Show hint
Look for a moment where the product decides for you instead of asking you.
Show answer
Model answer: A smart thermostat that shifts your schedule on its own. A leading signal would be how often it overrides your manual adjustment back to its own guess within an hour. A rising override rate would predict a wave of people getting annoyed and turning the whole thing off, weeks before anyone files a complaint.
Short answer, the number question
6. If the on-call engineer had raised the confidence threshold instead of lowering it, so Wend asked twice as often, would completion rate have been safe that week? Why or why not?
Show hint
Think about what a confirmation actually protects against, and what it costs.
Show answer
Model answer: Safer, but not automatically fine. Asking more often catches more of the shaky bookings before they ship, which protects completion rate. It also means more bookings sit waiting on a person to answer, and some get abandoned if nobody answers in time. The right call still needs the audit to check whether the extra asks were catching real problems, not just slowing everything down for no reason.
Before you close this one
Why this works
Tests whether you'll trust a metric that's tuned by the same lever it's supposed to be watching, or go find an independent check. Most candidates stop at naming a leading indicator and never ask who can quiet it.
Follow-up traps
"Isn't a rising confirmation rate just the agent being appropriately careful? Maybe that's fine on its own." Response: it can be. The problem wasn't the rise, it was what happened in week seven: nobody investigated it, someone just suppressed it. The audit is what tells you which one is happening.

"Why not just make Wend ask before every booking, then you never have a silent wrong action at all?" Response: that kills the product. A booking that waits on a person to answer isn't a ninety-second autonomous booking anymore, and some of those asks never get answered at all, so you'd trade completion rate for completion rate in a worse direction.
If pressed
The confidence threshold isn't one global number. It's set per task type, room-substitution confidence is tracked separately from fare-class confidence, so one hotel partner's new data format can push just the room-substitution threshold without moving the aggregate at all. The guardrail that actually catches this is a golden set sliced the same way, by task type, not one aggregate audit number.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more