Describe a leading indicator specific to agent products.
Find the number that moves before Wend's real booking success rate does, and don't let a quiet week convince you the problem is fixed.
- Put Wend's confirmation request rate on the weekly review, next to completion rate, not behind it.Why: it moved from three percent to fourteen percent over six weeks while completion rate held near ninety one percent the whole time.
- Never let a falling confirmation rate count as a fix on its own.Why: the rate and the agent's real accuracy can move in opposite directions the moment someone touches the same setting that produces both.
- Run a same day audit against a fixed set of past bookings whenever the confirmation rate moves more than five points in a week, in either direction.Why: this is the one check that tells a real improvement apart from a quieted alarm.
- Slice a rising confirmation rate by task type and partner before changing any setting.Why: at Journeo the whole six week rise traced back to one new hotel partner's data, not the entire agent losing its footing.
- Log the agent's mid task replanning rate as a second signal, not the main one.Why: replanning also rises before a real slip, but it's noisier for an agent that's supposed to replan when a fare changes, so it backs up the confirmation rate instead of replacing it.
- Give one named person the power to freeze the confidence threshold the same day a floor gets crossed.Why: a number nobody can act on the same day is just a chart on a wall.
How to answer this, stage by stage
Nobody is grading whether you can name a metric. They are grading whether you can catch a healthy looking tile lying to you before a customer does. Seven moves get you there.
Let's learn
On a TV bolted to the wall in Journeo's ops bay, one tile updates every morning: the share of Wend's bookings that finished without anyone stepping in to fix them.
Wend is a chat agent inside Journeo's app. Tell it your dates, your budget, and where you're headed, and it searches flights and hotels, picks the best match, and books both, with nobody clicking a single button.
Before Wend could book on its own, a person still picked from a results list and confirmed every leg by hand. That took about eighteen minutes for a return flight with a hotel attached, most of it spent comparing fares that all looked about the same. Now Wend finishes the whole thing itself in under ninety seconds, for roughly nine bookings out of every ten it touches. The tile on the wall has read close to ninety one percent for months.
Here is the turn. That ninety one percent never moved. Not once, for six straight weeks. And that steadiness is exactly the problem, because underneath it, the rate Wend paused to ask a person before finishing a booking climbed from three percent to fourteen percent, and the completion number was about to fall for real, not because Wend got worse on the tile, but because the tile was never built to notice.
At its worst, this doesn't look like an outage. It looks like nothing. A traveler checks into the wrong room type, or gets an extra night nobody asked for, sorts it out at the front desk, and never opens a support chat, because as far as Wend's own record shows, nothing went wrong.
What I would leave alone: Wend's confirmation rate for small stuff, window seat or aisle, doesn't need a floor at all. Getting that wrong costs someone one annoyed email, not a wrecked trip. Only the asks that stand between the agent and money actually leaving someone's card are worth watching this closely.
The lesson: a number that goes quiet can mean the problem got fixed. It can also mean somebody stopped measuring it out loud. Those look exactly the same on a wall-mounted tile, and the only way to tell them apart is to check something the tile itself doesn't control.
Now here is the same thing as a story
The short version sits above. Read on for the Tuesday a customer's thank you note was the only warning anyone got.
Demola Ferran has led product for Wend for two years, since before it could book anything on its own. Back then it only searched, and a person still clicked confirm on every leg by hand.
For most of a year after Wend started booking on its own, the tile in the ops bay barely moved. Ninety, ninety one, ninety two percent, back to ninety one. A boring number was the whole point. It meant Wend was doing today what it did yesterday.
The confirmation rate crept up so slowly nobody called a meeting about it. Three percent in week one. Four in week three. Six in week four. By week six it sat at fourteen, and the completion tile still read ninety one, steady as ever, because the trouble hadn't reached a real trip yet.
Then came a Friday afternoon in week seven. An engineer on call that week, watching the confirmation rate graph climb on a shared screen, did the fastest thing available to make the number stop climbing: turned the confidence threshold down a few points, so Wend would ask less often. The graph flattened out. He logged it as handled and went home for the weekend.
The completion tile never moved. It couldn't. It was reading what Wend said about itself, and Wend, asking fewer questions now, said done more often than not.
Two weeks later, Iona Corry opened a support chat from a hotel lobby in Lisbon, mid-trip with her family. Not to complain. To say thanks. Wend had booked her a room a size smaller than what she'd asked for, and the front desk had sorted it out at check-in with no trouble. "No big deal," she wrote. "Your app usually asks before it does anything like that though. This one it just went ahead."
Nobody flagged it. A five-star chat doesn't get read twice. It sat in a queue until a teammate, pulling transcripts for an unrelated project, noticed the line about Wend usually asking first and thought that was an odd thing for a happy customer to mention.
Demola pulled the real numbers that afternoon. The confirmation rate had been sitting near four percent for two weeks, which read as a win to anyone glancing at the tile. The post-trip audit, which only ran once a month, hadn't caught up to any of it yet. When it finally did, three weeks after the Friday the threshold got turned down, the real completion rate for that stretch came back at seventy eight percent.
Nobody at Journeo could say how many other trips like Iona's had gone quietly wrong in that window, because none of them had generated a ticket. A happy guest at a front desk leaves no record built to catch this.
Here is the part that actually mattered. Demola's team never had a real number for what happened on a trip in real time. They had a flag Wend set about itself, and a flag can be told to say anything the thing measuring it wants it to say.
Setting up that alert eighteen months earlier had taken about ten minutes in a stand-up. One number, one alarm, one threshold. It was the right call. Nobody could have known that the fastest way to quiet a rising alarm would be to turn down the very knob that produced it, and hand that knob to whoever happened to be on call that Friday.
Run the same six weeks through the fixed design. The audit runs same day now, any time the confirmation rate moves more than five points in a week. Week seven's drop gets flagged within hours, not caught by a customer's thank you note three weeks later. The threshold change gets reverted before Iona ever books her trip, and the real completion rate never leaves ninety one.
One design let the alarm get turned down by the same hand that set it off. The other makes that exact move show up as its own event, checked before it's allowed to count as fixed.
What I'd tell myself, back in that ten-minute stand-up: the day you build an alarm, ask who has the power to make it quiet, and whether quiet and fixed will ever look different to them.
LEAD, and the tile that couldn't tell fixed from quiet
This is a metric question, so LEAD does the work: find the outcome, find what moves first, name how it gets gamed, then say what you'd actually do.
Three things worth naming directly, since this is where the real judgment sits. We considered leading with the agent's own mid-task replanning rate instead of the confirmation rate, and set it aside as the lead metric: Wend is supposed to replan when a fare changes mid-search, so a rising replan rate is often healthy, and it's a harder number for anyone outside engineering to look at on a Tuesday and act on. We kept it running as a second, corroborating check instead. The failure worth naming by name is a form of confident wrongness: the model's own confidence calibration shifted after one hotel partner changed the shape of its availability data, so Wend started treating genuinely uncertain room-type matches as safe to book alone. The guardrail is a task-type-sliced golden set, a batch of past bookings with known right answers, checked against the confirmation rate every week rather than trusting the tile by itself. None of this is free. Asking more often is slower, and it can lose the booking entirely if nobody answers in time, which trades against the whole reason Wend exists, a flight and a hotel booked in under ninety seconds with nobody clicking anything. The bar we set is a calibrated one, not a hard rule: the confirmation rate and the golden-set wrong-action rate should move together, most weeks, within their normal band. When they stop moving together is the signal, not any single number on its own.
And if you want to be sure it really works, try it somewhere else
Same four letters, a warehouse reorder agent instead of a travel agent, and the trap is the same one dressed differently.
Northbank Supply runs Tally, an agent that watches stock levels across its warehouses and places reorders with vendors on its own, inside a spending cap, no purchase order needing a person's signature. Nkem Segovia runs procurement operations there.
L, link. Not whether Tally's own log says order placed. Whether the order that actually arrived matched what was needed, right vendor, right quantity, right price, checked against the receiving dock's own count, not Tally's own record.
E, early signal. The rate Tally pauses to ask a person to confirm an order before it submits one, on purchases it used to place alone under its spending cap. It rose from two percent in week one to eleven percent by week four, while the dock's error count stayed low the whole time.
A, abuse. In week five, procurement raised Tally's auto-approve spending cap so it would stop asking as often. The confirmation rate fell to two percent by week six. The dock's own count of wrong orders kept rising the entire time, from one wrong order a week to six.
D, decision. Same shape of table. Inside the normal band with the weekly dock audit clean, leave it. Rising for two weeks straight, slice by vendor before touching the cap. Any fast drop right after someone raises the cap: unproven until that week's dock count clears it.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip to the fix: whatever leading signal your dashboard shows, pair every reading of it with an independent check that doesn't share a setting with the signal itself.
Cost: engineering says a same-day audit is a month out. Don't read the live tile alone as a stopgap in the meantime, pull a hand sample of last week's bookings once a day until the real pipeline ships.
The model got better, for real: say Wend's underlying model genuinely improves next quarter. That still isn't the same claim as the confirmation rate falling for a good reason. A better model can lower that rate honestly, which is exactly why the number needs a second check that doesn't just take its word for it.
Where people run it wrong.
They watch the leading number and the lagging number on separate dashboards, so nobody ever puts them side by side and asks why one moved without the other.
They let whoever can change the underlying setting also be the one who reads the alert it produces.
They wait for a monthly audit to confirm what a weekly one would have caught in days.
How to use it live. Say the reframe before naming a single metric: "a number that used to predict trouble can be tuned to stop predicting it, without the thing it was watching ever getting better." That buys you room to give the real answer instead of reciting "track leading indicators" on reflex.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Why not just make Wend ask before every booking, then you never have a silent wrong action at all?" Response: that kills the product. A booking that waits on a person to answer isn't a ninety-second autonomous booking anymore, and some of those asks never get answered at all, so you'd trade completion rate for completion rate in a worse direction.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Leading vs lagging indicators for AI
- #1 Give three leading indicators of AI feature health and the lagging metric each predicts.
- #2 Why do lagging metrics fail you specifically in AI products?
- #3 Describe the leading indicators you would watch in the first 48 hours after an AI launch.
- #4 Explain how retry rate functions as a leading indicator.
- #5 What early signal predicts churn from an AI feature?
- #6 How do you build an early warning system for silent quality degradation?