How do session length and AI feature quality relate, and in which direction?
Session length and quality don't move in one fixed direction. The direction flips depending on what's actually happening inside the extra minutes, and you only find out by checking whether the session finished with the thing the user came for.
- Track completion inside each session-length band, never average length alone.Why: two sessions of the same length can have opposite meanings, one good and one bad, and only completion tells them apart.
- Rule out a tracking change before you trust any rise in the number.Why: a longer idle-timeout window or a new heartbeat ping makes sessions look longer with nobody's behavior changing at all.
- Recut the average by what the extra time was actually spent on.Why: a healthy average can be built from two opposite groups, one genuinely working, one stuck, and averaging them hides both.
- Name the real cause candidates before picking one.Why: confusion, a deeper task, and a chattier interface all raise the same number for completely different reasons.
- Run the one check that tells good long from bad long apart: did the session end in the outcome, not just more time in it.Why: time spent is what the model measures. Task finished is what the person came for.
- Route a flag to a person when long, low-completion sessions start clustering.Why: by the time it shows up as churn, the friction has already cost weeks nobody was watching.
Working the answer out loud, step by step
Nobody gets marked for knowing that "it depends." You get marked for showing exactly which number proves which direction, in under two minutes.
Let's learn
What does it mean when the number everyone's watching goes up, and nobody can say if that's good news?
Draftloom is a website a student opens with an essay draft. Paste it in, and it hands back comments on each paragraph: thin evidence here, a stronger transition there, a claim that needs a source. The student reads the comments, edits the draft, and checks it again if they want.
Before Draftloom added a chat feature, a typical session was short and single-pass: paste, read the comment list, apply what makes sense, done. The average session ran about six minutes, and 71 out of every 100 sessions ended with the student saving a version of the draft that was measurably different from the one they pasted in, meaning they'd actually used the feedback.
Then Draftloom shipped a chat box next to each comment: "explain this suggestion." A student could ask the model to say more about any single note instead of guessing what it meant. Average session length climbed steadily for six weeks, six minutes to fourteen. The team's weekly review called it out as a win. People were spending twice as long in the tool.
Here is the turn. Say it plainly: the extra time itself was never the problem. What changed was what the extra time was made of. For six weeks, more time in the chat correlated with nothing bad, because most of it was students genuinely working through a comment they didn't understand the first time. Somewhere in week six or seven, a second pattern started showing up inside the same average: a student would ask "explain this suggestion," get an answer, ask again in slightly different words, get almost the same answer, and never actually touch the draft.
At its worst, this doesn't look like a crisis. It looks like a metric quietly telling the opposite of the truth. Kiran, a ninth grader, opens Draftloom the night before an essay is due, gets a comment on a weak topic sentence, and asks the chat to explain it. The explanation uses a word Kiran doesn't fully follow, so Kiran asks again. And again. Eleven minutes pass. Kiran closes the tab without saving a single edit. On the dashboard, that session reads as a strong one: eleven minutes, well above average, exactly the kind of number the team had just started celebrating.
What I would leave alone: not every long session needs this check. Draftloom also has a library where students browse example essays with no draft of their own to finish. Time spent there was never supposed to end in a completion event, so a rising average in that section is not a warning sign. Splitting a number that has no outcome to check against just adds noise.
The lesson: time on a tool and value from a tool point the same way right up until the tool starts talking back. Once it can hold a conversation, a rising number can mean someone is finally getting through to the hard part, or it can mean they're stuck in a loop the model can't see either.
Now here is the same thing as a story
The short version sits above. Read on for the Thursday review where "engagement is up" almost became the whole headline.
Nasir Kettner has run product analytics at Draftloom for two years. He built the weekly review deck himself, and he's the kind of person who reads a number twice before he repeats it in a meeting.
For most of a year, the deck barely changed shape. Average session length sat near six minutes. Draft completion sat near seventy percent. Boring numbers, on purpose. Boring meant Draftloom was doing this week what it did last week.
The chat feature shipped on a Monday in March, and for the first two weeks Nasir did what he always did: he pulled the completion rate split by session length, the way the deck had always shown it. Both looked fine. Then, in week three, the team switched to a new dashboard tool that made the split view a second click instead of the front page. Nasir told himself he'd check it "when something looked off." Nothing looked off. The headline number, average session length, kept climbing, and climbing looked like winning.
By week six, Nasir had stopped opening that second click most weeks. Not a decision, exactly. Just a habit that quietly stopped happening.
The trigger was small. A support lead forwarded one message from a parent: "my daughter spent twenty minutes on this last night and turned in the same draft she started with, is that normal?" Nasir almost filed it as one anecdote. Then he remembered a second one from the week before, worded almost the same way.
He pulled the split that afternoon. Sessions where a student asked the explain feature the same question, in different words, more than twice, were finishing at 31 percent. Those sessions had been getting longer every week, right alongside the genuinely productive ones, and the one number on the front page of the deck had no way to tell them apart.
Here is the part that actually cost something. Draftloom's growth team had used the six-week climb in session length to justify pushing the explain feature to every grade level, including younger students who needed it least and got confused by it most. By the time Nasir's split reached the room, the feature had been live for three more weeks than it needed to be at that reach, and nobody could say exactly how many students like Kiran had opened Draftloom the night before something was due and closed it with nothing saved.
Three months earlier, in a fifteen-minute meeting, the decision to lead the weekly review with average session length felt obvious. It was the single number that had tracked completion almost perfectly for a year. Nobody put a date on that assumption. Nobody came back to ask if a new feature had just broken the link between the two.
Run the same six weeks through the fixed dashboard. The completion-by-band split sits on the front page next to the headline number, not a click away. In week four, the re-asking segment's completion rate crosses below its own floor while the overall average is still climbing and still looks fine. The gap gets flagged that Friday. Someone checks the transcripts, sees the same explanation phrased three ways for the same comment, and the team ships a small fix: after a second "explain this" on the same suggestion, the model asks a different question instead of repeating itself. Five weeks of unflagged confusion become one.
One dashboard watched a number that used to mean one thing. The other watches whether it still does.
What I'd tell myself, back in that fifteen-minute meeting: the day a single number earns the front page of a deck, write down what has to stay true for it to keep meaning what you think it means. Someday something will change underneath it, and the deck won't tell you when.
TRACE, and the five checks the front-page number skipped
This reads like a metric question, but the honest job here is finding out which of two opposite stories a single rising number is hiding, so TRACE does the work, not a straight measurement framework.
Three things worth naming directly, since this is where the real judgment sits. The rejected alternative was capping session length outright, forcing the chat closed after two or three turns. That was ruled out on purpose: it would have cut off the genuinely-revising group at 81 percent completion just as hard as the confused group at 31 percent, punishing the exact behavior the feature was supposed to reward. The AI-specific failure worth naming by name is a kind of silent degradation that has nothing to do with the model getting anything factually wrong: the explain feature answers every question correctly and still produces a worse outcome, because a technically right answer phrased three different ways doesn't teach a confused student anything new. The guardrail is a segment-level golden set, a batch of past chat threads labeled by whether the student's actual edit changed after the explanation, gating any change to the explain feature before it ships, not just checking whether the model's answers are accurate in isolation. And there's a real trade-off in the fix itself: making the model ask a different question after a repeated "explain this" instead of repeating the same explanation adds a small amount of extra latency and inference cost per turn, because the model now has to reason about what it already said instead of answering fresh each time. That's the trade Draftloom took: a slightly slower, slightly more expensive second turn, in exchange for catching a confused student in week four instead of losing her by week nine.
And if you want to be sure it really works, try it somewhere else
Same five letters, a field-service AI instead of a homework tool, and the split runs along the same fault line: is the extra time work, or is it the tool being confusing.
Coilwright makes an AI assistant for HVAC technicians. A technician opens it on a job, describes symptoms, and the assistant walks through likely causes and fixes. Lachlan Sarraf runs operations analytics there.
T, timeline. Average time-per-call with the assistant open rose from nine minutes to nineteen over five weeks, right after Coilwright added a feature that lets the assistant ask a technician clarifying questions instead of guessing from one description. First-visit fix rate looked steady for three weeks before it started to slide.
R, recut. Calls where the assistant asked two or three clarifying questions and then named a specific part: fixed on the first visit 84 percent of the time. Calls where the technician re-described the same symptom because the assistant's question didn't make sense to them: fixed on the first visit 38 percent of the time. Same rising average call time, opposite outcomes.
A, assume nothing. Coilwright checked first whether a firmware update on the handheld device had changed how idle time between taps gets counted. It hadn't. The rise in time-per-call was real technician behavior, not a counting artifact.
C, cause candidates. A technician working through a genuinely unusual fault. A technician confused by a clarifying question phrased in engineer language instead of trade language. A UI change that added a confirmation tap per step without changing what got fixed.
E, evidence test. Did the call end with a part named and a fix logged, not just more minutes on the clock. That's the one number that told a hard job apart from a confusing assistant.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the fix, whatever your rising number is, check completion inside each band before calling the rise good news.
Cost: the segment-level split needs new tagging that engineering says is a month out. Pull a manual sample of fifty long sessions by hand in the meantime, rather than trusting the blended average while you wait.
The model got better, for real: say the explain feature's underlying model genuinely improves next quarter, fewer wrong facts, better phrasing. That still isn't the same claim as fewer confused loops. A better model can even make the loop harder to catch, because each individual answer looks more convincing while the student still never gets unstuck.
Where people run it wrong.
They see the number rise, call it engagement, and move straight to rolling the feature out wider before checking what's inside it.
They build the completion split, then leave it as a report nobody's assigned to check weekly, instead of a number with a floor and an owner.
They fix a confusing loop with a training document telling users to "ask better questions," instead of a design change that catches the loop itself.
How to use it live. Open with the flip in one line: "a rising session length can mean two opposite things once a tool can talk back, and you can't tell which from the average alone." That buys room to give the real diagnosis instead of reciting "measure engagement" on reflex.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Isn't tagging every follow-up question to the suggestion it's about just extra engineering for a metric?" Response: it's the one piece of instrumentation that turns "sessions got longer" from a guess into a checkable claim, and without it the team has no way to know which half of the average it's celebrating.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Leading vs lagging indicators for AI
- #1 Give three leading indicators of AI feature health and the lagging metric each predicts.
- #2 Why do lagging metrics fail you specifically in AI products?
- #3 Describe the leading indicators you would watch in the first 48 hours after an AI launch.
- #4 Explain how retry rate functions as a leading indicator.
- #5 What early signal predicts churn from an AI feature?
- #6 How do you build an early warning system for silent quality degradation?