What does an increase in average conversation turns tell you? Give two opposite readings.
Read a climbing average conversation-turn count as two competing stories, not one verdict, before your next ops review votes on it.
- Split every session's turns into needed steps and re-check steps before you decide the number is good or bad.Why: a 79 percent jump can hide a rollout that's working and a trust problem that's spreading, and one blended line can't tell you which.
- Check a rise in needed steps against your own test set, not against a feeling.Why: more steps only means progress if the flow is still passing its own golden set most of the time.
- Trace a rise in re-check steps back to the one answer that started it.Why: nobody starts doubting a chatbot for no reason. It starts with one bad turn they remember.
- Give any re-check spike a named owner within the week.Why: a quiet trust problem spreads by word of mouth long before it turns into a ticket you can count.
- Weigh the extra tokens and extra minutes a harder flow costs against what it replaces.Why: walking a whole request end to end only earns its keep if it beats the old queue by enough to matter.
How to answer this, stage by stage
Nobody is grading whether you can name that a metric is ambiguous. They're grading whether you can build the one number that untangles it. Seven moves get you there.
Let's learn
Here's a strange kind of number. It can climb because a tool finally earned real work. It can climb because nobody trusts it at all. Same line, on the same chart. Opposite reasons.
Ravensworth Manufacturing built DeskPilot, a chat window every employee opens for the boring stuff: reset a password, ask for software access, get help when something's broken. Type a question, get an answer, keep working.
For most of the year, an average DeskPilot chat ran 2.4 turns. Ask, answer, maybe one follow-up, done. That number barely moved.
Then, over six weeks in the spring, it climbed to 4.3 turns, a jump of 79 percent. Ilkay Denholt, who owns the dashboard, watched it happen in real time and had one honest reaction: something is different here, and I don't yet know if it's the good kind of different.
Here is the turn. A rising average is not a mistake count. It's a shape two very different mornings can both draw. One employee brought DeskPilot a harder job than it ever used to get, and the bot walked her through it: seven small turns instead of one. Another employee stopped believing DeskPilot's first answer, and now checks every claim before she'll act on it: also seven turns, for the opposite reason. Both mornings land on the same number on Ilkay's chart.
At its worst, this costs two things at once, and they pull in opposite directions. Miss the good reading, and you cancel a rollout that was working, right as it started paying off. Miss the bad reading, and a trust problem that started with one wrong answer keeps quietly costing people their afternoons, invisible on a dashboard that only ever shows one blended line.
What I would leave alone: a plain lookup, like the guest wifi password, should never need more than one turn, and if that one starts climbing, that's worth a look, good story or bad. The troubleshooting flow is different. It's built to ask three or four questions on purpose before it gives a fix, so a high turn count there was never a warning sign, and chasing it would just waste everyone's week.
The lesson: an average turn count is not a health score. It's a blend of every reason a person might need to keep talking to a bot, good reasons and bad ones stirred into the same cup. The day you stop asking what's in the blend is the day the blend starts deciding things for you.
Now here is the same thing as a story
The short version sits above. Read on for the Wednesday morning that almost got a working rollout cancelled by its own success.
For a year, DeskPilot ran Maren Sundahl's simple rule: quick stuff goes to the bot, anything that touches more than one system goes to a person. This spring, without anyone announcing it, that rule quietly stopped being true.
Maren schedules the floor at Ravensworth's Dayton plant. Nine years in, and she can tell you which of the plant's three approval systems will be the slow one before she's finished her coffee.
When DeskPilot launched, it was fast and it was honest about its limits. Ask it to reset a password, done in one turn. Ask it for CAD license access, and it said plainly: I can't walk you through that, here's the ticket form. Maren tried the harder asks twice, early on, got the same honest bounce both times, and learned the boundary. After that she didn't bother. Quick stuff to the bot. Everything else, straight to the queue, same as everyone on her shift.
The boundary moved before she ever noticed it moving. First, DeskPilot started asking one more question after a bounce, instead of pointing straight to the ticket form. Then it started checking a manager's approval status itself, instead of asking her to chase it down. Then, on a Tuesday nobody made an announcement about, the bounce just stopped happening.
The trigger was small. Her line lead mentioned, almost in passing, that he'd gotten his own badge access sorted "in like five minutes, through the chat thing." Maren filed that away and didn't think about it again until she needed it.
So on a Wednesday morning, with a rush tooling order stuck behind a CAD license she didn't have, she typed the request into DeskPilot instead of the ticket form. Seven turns. Confirm who she was. Confirm which system. Confirm which license. A ping to her manager's phone. A short wait. A confirmation. Access granted. Six minutes, start to finish. The ticket route she'd have taken the week before averaged two and a half business days.
Maren never had a number in her head either. She had a boundary: things the bot could do, and things it couldn't. The day the boundary moved, she didn't ease into the harder asks. She just started bringing them, the same way she'd bring anything to a person she'd suddenly learned to trust with more.
A year earlier, in the room where DeskPilot's first dashboard got built, the decision took about ten minutes. One number, average turns per session, and a rule that said if it climbs, look into it. Nobody argued. The bot only did one kind of thing back then, so one number described it fine.
Run that same Wednesday through the fixed dashboard. The aggregate still climbs 79 percent. But now it's two lines, not one. Needed steps up 76 percent, holding steady against a golden set still clearing 91 percent. Re-check steps essentially flat for access requests, the kind like Maren's. The ops review doesn't blink at the total anymore, because the total finally has to show its work. The access-request rollout ships on the Friday it was scheduled for.
What I'd tell myself, back in that ten-minute meeting: we built a health check for a bot that only ever did one thing, and we never came back to ask whether that was still true. It wasn't, for months, before anyone noticed.
Run FLIPS twice before you trust one climbing line
Same five letters both times. Only the middle one, I, actually changes, and that's the whole trick to keeping two true stories straight on one chart.
I, run the other direction. Trusts the first answer and moves on, versus cross-examines every claim before acting on it. This is the verification flip, the same family a rising number always tempts a candidate toward first. It's the honest answer for Robyn's Tuesday, not for Maren's.
P, a different old decision. DeskPilot answered a compliance-flagged setting with the same confidence it uses for anything else, no stricter bar for the one category where being confidently wrong actually costs someone their afternoon.
S, a different replay. The doubling in re-check steps gets traced, within two days instead of a full quarter, back to the one wrong instruction that started it. Fixed, and Robyn's sessions settle back near two turns inside two weeks.
Three things worth being straight about here. We looked at a hard ceiling on turns per session, forcing a handoff past five, and turned it down: it can't tell a task that finally needs more room from a person who's stopped trusting the room it already has. The failure worth naming by name is a confidently wrong answer on a compliance-flagged setting, the kind of mistake that reads as helpful right up until it isn't. The safety rule for it is a stricter bar for exactly that category: a current, cited source, or a handoff to a person, checked against its own smaller golden set. And the trade being accepted is real. An access-request session now costs roughly three times the tokens a password reset does, and six extra minutes of someone's morning, against a queue that used to cost two and a half days. Worth it here. Not worth building into the one-line lookups, which is exactly why we split two lines and not twenty.
And if you want to be sure it really works, try it somewhere else
Same question, a repair van instead of an office chat window, and the third reading has nothing to do with trust or task difficulty at all.
Kilbrenner Field Services runs FieldFix, a chat tool techs open on their phone to look up a fix for a fault code on a furnace or an AC unit. Talwinder Ostvald has driven a Kilbrenner van for six years and knows most of the plant's common faults on sight.
For FieldFix's first few weeks, a typical troubleshooting session ran about 2 turns: describe the fault in one message, get a fix. Over the next five weeks, the average climbed to 5 turns, up 150 percent. Not because the faults got harder. Talwinder was seeing the same routine short-cycling and no-heat calls he'd always seen. And not because trust broke. His first-fix rate held steady around 88 percent the whole time.
Here's what actually happened. FieldFix badly parsed long, rambling fault descriptions, the kind Talwinder gave it naturally, and when it choked it just asked a vague, generic "can you clarify that." No hint about which part it missed. So Talwinder learned, the slow way, through five weeks of trial and error, that short, separate messages worked better: the error code first, the unit model second, the symptom last. His turns climbed because he was learning to speak the bot's dialect, not because anything about the job or his trust in it had changed.
The old decision FieldFix's team would take back: it accepted free text with no receipt showing what it caught and what it missed, just a generic nudge to clarify. The fix is one line after every message, showing exactly what got captured and what didn't, so a tech sees the gap on message one instead of learning FieldFix's shape over five weeks of guessing. Replayed with the receipt, Talwinder's first message gets an instant reply: got unit 12, short cycling, missing error code. He adds the code in one more turn. Two turns, not five, and nobody spends a month learning a dialect nobody ever wrote down.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the fix, whatever the average shows, split it into steps the task needs and steps spent re-checking, before you call the climb good or bad.
Cost: engineering says a real step-type split is a quarter out. Don't read the raw average alone as a stopgap. Pull a rough split by hand from a sample of transcripts every week until the real one ships.
The model got better, for real: say completion rates genuinely climb next quarter, across every flow. That still isn't proof the extra turns are healthy. A better model can make a trust problem easier to miss, because the good story's climb covers for the bad one hiding right behind it in the same average.
Where people run it wrong.
They build the split, then leave it as a chart nobody owns, instead of a number with a floor and a name attached.
They wait for the quarterly review to catch a trust problem that's been running for weeks, instead of a same-week check that would have caught it in days.
They answer a scary aggregate with a plan to add more review, instead of a second line that tells them whether review was ever the actual problem.
How to use it live. Say the reframe before naming either story: "a rising average can mean the tool just earned harder work, or that nobody trusts one answer anymore, and it looks identical from the chart alone." That buys you room to actually answer the question, instead of guessing which reading the interviewer wants to hear.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Isn't splitting task-needed and re-check steps just relabeling the same average after the fact?" Response: no. Each flow's steps get labeled needed by design, before a single session runs. Identity check, approval ping, and provisioning are known steps in the access-request flow. Nobody's picking the split after seeing which story it flatters.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Leading vs lagging indicators for AI
- #1 Give three leading indicators of AI feature health and the lagging metric each predicts.
- #2 Why do lagging metrics fail you specifically in AI products?
- #3 Describe the leading indicators you would watch in the first 48 hours after an AI launch.
- #4 Explain how retry rate functions as a leading indicator.
- #5 What early signal predicts churn from an AI feature?
- #6 How do you build an early warning system for silent quality degradation?