ConceptFoundationalQuality, Cost & Token Economics / Leading vs lagging indicators for AI / #1
Give three leading indicators of AI feature health and the lagging metric each predicts.
Three numbers move before an agent ever stops opening the panel. Here is what each one is actually warning you about, and where the same alarm gets rung by accident.
4% → 21%
Override rate on high-confidence notes, six weeks
6s → 35s
Gap: promised handle time vs. real handle time
81% → 46%
Notes actually opened live, not just present
The direct answer
Do not wait for the panel-use chart to move. Watch three earlier signals instead: how often agents wave off a note the model marked as sure, how far the real call time drifts from what the note promised, and how often agents open the note while the call is still live rather than never. Each one moves weeks before panel use, complaint counts, or the trust survey ever budge, and each pairs to one lagging number you can name in advance.
Do this, in order
Watch the override rate on high-confidence notes first.Why: it moves earliest, and it catches distrust before agents stop opening the panel at all.
Pair it with the gap between the note's promised handle time and the real one.Why: this is what predicts a supervisor's inbox filling with complaints, not the model's own offline accuracy score.
Track live-reference rate, never panel-open rate.Why: a panel that opens by default looks used on every single call whether or not anyone reads it.
Set a threshold that triggers a hand review, not a color change on a dashboard.Why: 6 percent override is normal caution. 12 percent sustained is a real problem. Only a person pulling 20 calls can tell the two apart.
Never grade an agent on the time gap while a segment is under review.Why: penalizing a person for a number the model caused breaks trust twice over.
Gate every new coaching prompt or model version behind a golden set before it reaches a live call.Why: this is the one step that stops the whole cycle from repeating on the next new call type.
How to answer this, stage by stage
The interviewer is not grading whether you can name three metrics. They are grading whether each one earns its place with a real lagging number attached, and whether you know what to do the day it crosses a line. Eight moves get you there.
1
Anchor it to one contact center, not "AI" in general
Say it like this
"Let me ground this. Coldharbor Roadside Assistance runs Signalwell, a tool that scores live calls and drops a coaching note into a side panel while the agent is still talking. Nadiya Boruc runs the coaching program there, about a hundred and forty agents across three shifts."
Why this works
A metric question answered in the abstract is unfalsifiable. Naming the product and the person makes every number you say next checkable.
2
Say what you're actually optimizing for, not the vanity number
Say it like this
"We considered just watching how often the panel was open. We ruled that out. It auto-expands at the start of every call, so it can read ninety percent for months while agents have quietly stopped reading a word of it. What actually matters is whether agents are still leaning on the note on their own, six months in, because that's the only thing that moves handle time and callbacks for real."
Why this works
Names the rejected alternative early and states the real business outcome (the L step) before naming a single leading number.
3
Name the three signals, each tied to its own future number
Say it like this
"One, override rate on notes the model marks high confidence, that predicts a drop in weekly active use about four to six weeks out. Two, the gap between the note's promised handle time and the real one, that predicts a rise in supervisor complaints about two to three weeks out. Three, how often agents open the note while the call is still live, not after, that predicts a drop in the quarterly trust survey score."
Why this works
This is the E step, and it's the actual answer. Every leading indicator gets a name, a timeframe, and the specific lagging metric it predicts.
4
Say exactly how each one lies to you
Say it like this
"Each one can be gamed or misread. Loosen what counts as 'high confidence' and the override rate drops without a single agent trusting the tool more. A rising time gap can just mean a genuinely harder batch of calls, not distrust, so you check call complexity before you act. And live-reference rate can fall because someone finally memorized the pattern, not because they stopped trusting it, so you check whether call quality held before you call it a real drop."
Why this works
The A step. Naming how a metric gets gamed is what separates a candidate who's built a dashboard from one who's actually watched people route around it.
5
Give the thresholds, not just the direction
Say it like this
"Under six percent override, do nothing, that's normal caution. Six to twelve percent sustained two weeks, pull twenty overridden calls by hand and check if it's a new call type the model hasn't seen much of. Above twelve percent sustained, freeze the coaching note for that call type and fall back to a plain checklist until a fixed version clears review."
Why this works
The D step. A metric with no threshold is a chart nobody acts on. Thresholds are set as a bar the model clears most of the time, not a promise it will never miss.
6
Prove it with the eight weeks nobody caught in time
Say it like this
"Here's what happens without this. Coldharbor shipped a prompt update that quietly dropped the caveat line on lower-confidence guesses. Override rate climbed from four percent to twenty-one percent over six weeks. Panel-use, the number leadership was actually watching, stayed at ninety-one percent the whole time, because the panel still auto-opened on every call. Nobody looked twice until repeat callbacks on one call type nearly tripled."
Why this works
A compressed, four-sentence failure does more work here than a paragraph explaining the theory would.
7
State the guardrail that stops it happening twice
Say it like this
"The real failure was that the model's 'sure' tag was wrong. A brand-new towing procedure got marked high confidence when the model had barely seen it. The fix is a golden set that always includes fresh examples from any new call type, and a new prompt or model version only ships once it agrees with hand-graded specialists on that set, most of the time, before it ever reaches a live call."
Why this works
Names the AI-specific failure mode (the model's own "sure" tag going quietly wrong on a kind of call it hasn't seen much of) and its guardrail in one breath, in threshold language rather than a promise of perfection.
8
Close on the trade you're accepting
Say it like this
"Putting the caveat sentence back into the note, and having it explain briefly why a guess is or isn't confident, adds about a second and a half to how long the note takes to generate mid-call, plus a bit more cost per call. I'd take that trade every time over agents quietly deciding, on their own, that the green flag can't be trusted."
Why this works
Naming the trade between quality, speed, and cost out loud is what makes this a real decision instead of a wish that all three were free.
If you remember one thing
A leading indicator is only worth naming if you can say the lagging number it predicts, the timeframe, and the threshold where you'd act. Anything short of that is a chart, not a metric.
Let's learn
What tells you a coaching tool is dying before the chart everyone watches ever moves?
Signalwell sits inside Coldharbor Roadside Assistance, an auto club's contact center. It listens to a live call, scores it, and drops a short coaching note into a side panel: what to say next, roughly how long the call should take, a flag for whether the model is sure or just guessing.
Three separate numbers, all answering the same real question: will the agent still be leaning on this thing in month six.
Before Signalwell, an agent handling a tow request worked from memory and a printed script taped to the monitor. The average call ran about six minutes, and a third of tricky ones needed a supervisor pulled in. In its first months, Signalwell cut that down close to four minutes, and pulled-in supervisor calls dropped by half. Agents liked it. Nadiya Boruc, who runs the coaching program, watched one number on her weekly report: the share of calls where the panel was open. It sat at ninety-one percent, and it stayed there.
Override rate on high-confidence ("green flag") notes, six weeks
A quiet prompt update in week 0 dropped the caveat line on lower-confidence guesses. Override rate crossed the 12 percent review line by week 3, three weeks before panel-use ever dipped.
Here is the turn. The rising override rate was not the real problem. It was the earliest honest signal Coldharbor had. What mattered was what caused it: a new towing procedure for electric vehicles had rolled out two weeks earlier, and Signalwell had barely seen it, but a bug in how the model judged its own certainty still tagged its guesses on those calls as "high confidence." Agents on the overnight shift, who handle most tow requests, got burned by a few wrong green flags and stopped trusting green flags generally, not just on EV calls.
We did not lose accuracy on paper. We lost the one thing the whole tool ran on: the agent's willingness to believe the green flag without checking it.
Knowledge spark: what is a live-reference rate?
The share of calls where the screen log shows an agent actually opened or expanded the coaching note while still on the call, not just that the panel was sitting there. A panel that auto-opens by default can look used on every call even when nobody reads a word of it.
At its worst, this cost more than an awkward number on a report. The gap between what Signalwell's note promised ("wrap this in about ninety seconds") and how long the call actually took widened from six seconds to thirty-five seconds over five weeks, on EV tow calls especially. Supervisors started fielding complaints about "the tool making calls take longer," rising from two a week to thirteen a week. And repeat callbacks on EV flatbed dispatch, where the customer had to call back because the first call went sideways, climbed from six percent to nineteen percent.
Weekly active panel use: the number leadership actually watched
Panel-use held flat at 91 percent for nine straight weeks after the problem started, because the panel opens by default. It only cratered once agents began closing it by hand faster than they opened it.
The choice I would take back is not the prompt update itself. It is that nobody put a check on it before it shipped. Coldharbor tested the new, friendlier wording on a general sample of calls and it read fine. Nobody built a slice of the test set from the newly onboarded EV procedure, because at the time of testing, almost no EV calls existed yet to sample from.
What I would leave alone: the panel-open auto-expand behavior itself. That default is genuinely fine, it just cannot be the metric anyone trusts, because "present" and "used" are not the same claim.
The lesson: a metric that is watching the panel instead of the person will read healthy for months after the person has already quit trusting what's on it. The gap between those two dates is exactly the gap the three earlier signals are built to close.
Now here is the same thing as a story
The short version sits above. Read on for the six weeks Nadiya spent explaining a number that hadn't moved yet.
Nadiya Boruc has run Coldharbor's coaching program for five years. She built the training deck every new agent gets on day one, and she can tell from the shape of a call transcript, before she checks a single score, whether an agent is having a good week or a bad one.
When Signalwell launched, the good months looked steady. Morning shift-change, she'd pull the weekly report, see panel-use sitting at ninety, ninety-one percent, and move on to the next thing on her list. The number never gave her a reason to look closer.
Underneath that flat number, something was thinning out. In week one after the EV procedure rolled out, an overnight agent got a green-flag note that told her the customer's flatbed request needed no special handling. It did. The customer's EV had a battery isolation switch the note never mentioned, and the call ran nine extra minutes while the agent figured it out live. She mentioned it to a coworker at the next break. By week three, three more agents had their own version of that story.
The number on the report and what the agent actually did with the tool can drift apart for weeks without anyone noticing which one is lying.
None of them filed a ticket. They just started reading the whole note more slowly, then skimming it, then, on calls that felt routine, not opening it at all. Nobody decided this out loud. It built the way distrust always builds: in private, one call at a time.
The moment anyone outside the overnight shift noticed was not dramatic. A regional manager forwarded Nadiya a spreadsheet: repeat callbacks on flatbed dispatch, up from six percent to nineteen percent over six weeks. Nadiya's first instinct was to check panel-use. Ninety-one percent. Healthy. She almost closed the spreadsheet.
The number she trusted was telling her the truth about the screen. It had nothing to say about what was happening in the agent's head.
It was only when she pulled call recordings, not dashboard numbers, that she heard it: agent after agent glancing at an open panel and doing the call from memory anyway. The panel was there. Nobody was reading it.
A year earlier, when Coldharbor built the coaching dashboard, watching panel-use made sense. Back then the panel didn't auto-open, so a low number meant exactly what it looked like: agents choosing not to use the tool. Nobody in that early design meeting pictured a future where the panel opened itself by default and the real signal moved somewhere the dashboard couldn't see.
Run the same six weeks through the fixed design. Override rate on green flags crosses twelve percent by week 3. The threshold fires a hand review that week, not week nine. A QA analyst pulls twenty overridden EV calls, finds the bug in an afternoon, and Signalwell's next prompt version ships with a golden set built specifically from the new procedure. Callback rate on flatbed dispatch never gets past eight percent. Nadiya spends her Tuesday on something else.
What I would tell myself, before any of this: the day a number stays flat by design, because of how the interface defaults, is the day to stop trusting that it means what it used to mean.
LEAD, the four moves behind this answer
This is a metric question, so LEAD runs it, not FLIPS. Each letter gets its own answer here, then the same four letters run again in Section 4 on a completely different product.
L
Link. The business outcome that actually matters.
Not the model's offline accuracy score. Whether agents at Coldharbor are still leaning on Signalwell's note on their own, six months in, because that's what actually moves handle time and callback rate.
Rejected alternative: watching raw panel-open rate. It stays high by default and hides real distrust for months.
E
Early signal. The thing that moves weeks before the outcome does.
Three of them, each paired to its own lagging metric: override rate on green flags predicts panel-use dropping, the handle-time gap predicts supervisor complaints rising, live-reference rate predicts the trust survey score falling.
This is the actual answer to the question. Everything else supports it.
A
Abuse. How the metric gets gamed or misread.
Loosen the confidence bar and override rate drops without real trust rising. A rising time gap can mean genuinely harder calls, not distrust. A falling live-reference rate can mean mastery, not distrust, if call quality held.
Each one needs a second check before you act on it, not a reflex.
D
Decision. What you'd actually do at each threshold.
Under 6 percent override, nothing. 6 to 12 percent sustained, hand review of 20 calls. Over 12 percent sustained, freeze that call type's coaching note and fall back to a plain checklist until a fixed version clears the golden set.
A metric with no threshold attached is a chart nobody acts on.
Two things worth naming directly, since this is where the real judgment lives. The AI-specific failure mode here is the model's own "sure" tag going quietly wrong on a kind of call it hasn't seen much of: a new call type the model has barely seen gets tagged high-confidence anyway, because its sense of "sure" was learned on the old mix of calls. The guardrail is a golden set that always carries fresh examples from any newly onboarded call type, and a rule that a new prompt or model version only ships once it agrees with hand-graded specialists on that set most of the time, checked before rollout, not after an agent gets burned. The trade-off worth stating plainly: putting the caveat sentence back and having the note briefly explain its own confidence adds real generation time mid-call, on the order of a second and a half, plus a bit more cost per call. That is a cost worth accepting, because the alternative is agents deciding on their own, silently, that green flags cannot be trusted, which costs far more than a second and a half ever will.
And if you want to be sure it really works, try it somewhere else
Same four letters, a grocery chain's self-checkout loss-prevention tool instead of a call center, and this time the leading signal is a dismiss rate instead of an override rate.
Corrales Grocery runs Loomstack at its self-checkout lanes. It watches the scanner feed and flags carts that likely have an unscanned item, sending an alert to a loss-prevention associate's radio. Zahra Pellerin runs loss prevention across the chain's twelve stores.
The break happens at step three, long before anyone can see it in the store's shrink number at the far end.
L, link. Not flag volume. Actual shrink dollars recovered per store per month, the number that decides whether Loomstack keeps its budget. E, early signal. The share of high-confidence flags an associate dismisses without walking over to open the cart. That predicts a rise in shrink dollars at that store, roughly three to four weeks later, once word gets around that flags don't need checking. A, abuse. A store manager can force the dismiss rate down by making associates walk over to every flag on principle, even ones they already know are false, which looks like trust and changes nothing about real theft. Misread the other way: dismiss rate can rise because a new store-brand barcode rollout spiked false positives, which is model drift, not associate distrust, so you check the false-positive rate before blaming the person. D, decision. Under 8 percent dismiss rate, healthy. 8 to 20 percent, spot-check 15 dismissed flags by hand. Above 20 percent sustained, pull that store's threshold back to human-reviewed samples only until the model is retaught what the new barcodes look like.
Same shape, different floor
A rate that stays low can mean real trust or a policy forcing the motion without the belief. Volume of compliance and honesty of compliance are not the same number, in a call center or on a store floor.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the fix: whatever adoption number you're watching, pair it with a share that shows whether the tool is actually being trusted, not just present.
Cost: engineering says a real live-reference metric can't ship for two months. Do not fall back on raw panel-open rate as a stopgap in the meantime, and hold off on any new alert built from that number until the real one exists.
The model got better, for real: say Signalwell's suggestions genuinely improve next quarter. That alone does not prove override rate falling means renewed trust, agents who already stopped checking may never notice the improvement. Only a rising live-reference rate proves people are actually engaging with the better output.
Where people run it wrong.
They treat "the panel is open" as "the note is being used," when it just means the interface auto-expanded.
They fix the number by loosening what counts as high confidence, instead of asking why agents stopped trusting the real ones.
They wait for the client's scorecard or a regional audit to notice, instead of checking these three signals on their own schedule.
How to use it live. Say the reframe before naming a single number: "A tool being open on someone's screen isn't the same as someone trusting what's on it, especially when it opens itself by default." That buys you room to give the real answer instead of reaching for a generic engagement metric.
Flashcards (click a card to flip it)
1 · THE FRAMEWORK
Which framework does this question use, and what's its one-line job?
Tap to flip
ANSWER
LEAD. Find the signal that moves first, weeks before the outcome the business actually cares about.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Nadiya Boruc, who has run Coldharbor Roadside Assistance's coaching program for five years, overseeing about 140 agents on Signalwell.
3 · THE HABIT
What did agents stop doing that used to work?
Tap to flip
ANSWER
Reading the coaching note in full before acting on it. They kept the panel open by default but stopped actually opening or reading it live, especially on calls that felt routine.
4 · THE THREE SIGNALS
Name the three leading indicators and what each one predicts.
Tap to flip
ANSWER
Override rate on green-flag notes predicts falling panel-use. The handle-time gap predicts rising supervisor complaints. Live-reference rate predicts a falling trust survey score.
5 · THE REJECTED METRIC
What metric did this answer rule out, and why?
Tap to flip
ANSWER
Raw panel-open rate. The panel auto-expands by default, so it can read healthy for months while agents have quietly stopped reading it.
6 · THE NUMBER
Fill in the blank: override rate climbed from 4 percent to ___ percent over six weeks, while panel-use stayed flat at ___ percent the entire time.
Tap to flip
ANSWER
21 percent; 91 percent. The gap between those two numbers is the whole point of watching a leading indicator at all.
7 · THE GUARDRAIL
What AI-specific failure caused this, and what stops it next time?
Tap to flip
ANSWER
The model's own "sure" tag went quietly wrong on a new call type (EV flatbed dispatch) it had almost never seen. The guardrail is a golden set that always includes fresh examples from any new call type before a version ships.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product. Which product, and what's the leading signal there?
Tap to flip
ANSWER
Loomstack, Corrales Grocery's self-checkout loss-prevention tool. Leading signal: the dismiss rate on high-confidence flags, predicting a rise in shrink dollars weeks later.
Check yourself Score: 0 / 0
Fill in the blank
1. Override rate on Signalwell's high-confidence notes climbed from 4 percent to ___ percent over six weeks, before panel-use ever moved.
Show hint
Check the first chart in "Let's learn."
Show answer
21 percent. It crossed the 12 percent review line by week 3, three weeks before the panel-use number showed any change at all.
Multiple choice
2. Why is live-reference rate a stronger signal here than raw panel-open rate?
A. It is cheaper to compute from the call logs.
B. Panel-open rate stays high by default even when nobody is reading the note.
C. It only works for calls under two minutes long.
D. It measures the model's own confidence instead of the agent's behavior.
Show hint
Think about what auto-expand does to a panel-open count.
Show answer
B. The panel auto-opens on every call, so "open" and "read" are not the same claim, and only one of them tells you whether trust is real.
True or false
3. True or false: a rising override rate always means the model has gotten worse.
True
False
Show hint
Check the "abuse" step. There's more than one honest cause for a rising override rate.
Show answer
False. It can also mean a genuinely new, undertrained call type, or a team quietly loosening what counts as "high confidence." The rate alone does not tell you which.
Multiple choice
4. What should happen once the handle-time gap stays above 25 seconds for three straight weeks?
A. Immediately lower every agent's performance score for that call type.
B. Trigger a speed and quality review, and pause scoring agents on that segment until it's fixed.
C. Delete the handle-time estimate from the note entirely.
D. Wait for the next quarterly trust survey to confirm it.
Show hint
Whatever caused the gap, it wasn't the agent's fault.
Show answer
B. Penalizing agents for a number the model caused breaks trust twice. Fix the cause first, and hold the scorecard while you do.
Short answer, apply it yourself
5. Think of an app or tool you use regularly. Name one leading signal it could watch, and the lagging outcome it would predict weeks in advance.
Show hint
Look for a small behavior that changes before someone actually quits using the thing.
Show answer
Model answer: A budgeting app. Leading signal: the share of auto-categorized transactions a user manually recategorizes each week. A rising share predicts, about a month out, that the person will stop opening the app's weekly summary at all, because they've stopped trusting the categories it shows them.
Short answer
6. If Coldharbor had waited for the quarterly trust survey instead of watching override rate, roughly how many weeks would have passed before anyone caught the problem?
Show hint
The prompt update shipped in week 0. The survey only runs once a quarter.
Show answer
About 13 weeks, versus about 3 weeks if override rate had been watched from the start. That ten-week gap is roughly the same stretch where callback rate on flatbed dispatch tripled unchecked.
Before you close the answer
Why this works
Tests whether you can name a metric that would look perfectly healthy right up until the morning everything broke, and explain why, instead of naming the model's own accuracy score and calling it a day.
Follow-up traps
"Couldn't you just lower the confidence bar for a 'green flag' so fewer notes ever count as high-confidence, and the override rate looks fine by definition?" Response: that's exactly the gaming case named in the A step, which is why the threshold review pulls actual calls by hand instead of trusting the rate alone.
"What if override rate rises because the model genuinely got better at flagging its own uncertainty, and agents are just disagreeing with correct low-confidence calls?" Response: check which confidence band the overrides cluster in. Overrides concentrated on notes marked high-confidence is the distrust signal. Overrides spread evenly across confidence levels is normal healthy skepticism, not a threshold trigger.
If pressed
The golden set isn't static. Coldharbor rebuilds it every time a new call type crosses roughly 200 live calls, because a set built before a procedure existed can't catch a "sure" tag going wrong on a procedure it never saw.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.