What product decisions explain why some AI note-takers retain users and others do not?
Notchwell and Loomjot both listen to a home visit and draft a clinical note afterward. Colm Whitfield is a hospice social worker who carries a caseload of fourteen families and has used both, one after the other, over the same year.
- Check whether the product keeps a searchable history of past notes at all.Why: this single decision is the biggest driver of the retention gap between the two.
- Look for when the two usage curves actually split, not just where they end up.Why: the split happened around week five, long before anyone's dashboard flagged a problem.
- Slice retention by role, not by an overall average.Why: case managers who need cross-visit history churned far faster than staff doing one-off intake notes.
- Rule out a model or transcription bug before blaming behavior.Why: transcription accuracy stayed flat the whole time; the drop was a habit change, not a quality change.
- Find the one action that predicts churn before it happens.Why: a failed "find my last note" search, right before someone stops opening the app, is the tell.
- Leave the one-off intake note flow alone.Why: a note nobody ever needs to look up again doesn't need a history feature to feel complete.
How to answer this, stage by stage
Six moves, because this question is narrower than it looks: it's really one diagnosis question wearing a comparison's clothes.
Let's learn
Every evening, Colm used to spend about twenty minutes after his last visit typing up notes by hand: what he observed, what the family asked for, what he'd flagged for the nurse. It was slow, but he could always flip back through his own notebook.
Loomjot arrived first, rolled out by his agency. It listened during the visit and had a clean note waiting by the time he got to his car. For a few weeks it was simply faster. Then he needed to check something he'd noted three visits back, about a fall risk at one family's home, and Loomjot had no way to show him. Each new note replaced the last one on screen. The old one wasn't gone from a server somewhere, but there was no way to find it, search it, or even know it existed.
Colm didn't complain. He started keeping his own paper tally of anything he'd want to reference later, on top of using Loomjot, which meant he was now doing two note-taking tasks instead of one.
At its worst: Colm nearly missed noting that a fall-risk flag from three visits ago had never been passed to the nurse, because Loomjot had no way to show him he'd raised it before, and he only caught it by checking his own paper tally.
What I would leave alone: a one-time intake note, for a client Colm will likely never revisit, doesn't need a searchable history at all. Building that feature for every single note type would be solving a problem some notes never have.
The lesson: a note-taker's value doesn't live in the moment it drafts a note. It lives in the moment, weeks later, someone needs to check what they said before. A product that only ever shows you today has quietly decided that moment doesn't matter.
Now here is the same thing as a story
The short version above is what you'd say defending this diagnosis live. Read this one for how the pattern actually surfaced.
Colm Whitfield has carried a hospice caseload for four years, and he can walk into a home he hasn't visited in three weeks and remember, within the first minute, exactly what he'd flagged last time.
Loomjot's agency rollout went smoothly for the first month. Then, in week five, Colm needed to check a note from a visit three weeks back and hit a dead end: no search, no archive, no way in.
He didn't file a complaint. He started a private paper tally alongside Loomjot, the workaround of someone with too much at stake to just hope he'd remember.
His agency's usage report didn't flag anything for months. Active-use percentage was still a healthy-looking number on the surface, and nobody had sliced it by role.
Three explanations were on the table once someone finally looked closely: staff had simply stopped using it (abandonment), staff had built their own private system around it (workaround), or staff had started keeping their spoken notes shorter and vaguer so there'd be less to forget (pre-editing).
The evidence test settled it: the churned cohort's failed-search rate climbed for two straight weeks before their last session, while their transcription-accuracy stayed exactly as good as everyone else's. The habit didn't fade because the notes got worse. It faded because the notes became unreachable.
Loomjot's team had decided, back at launch, that a note only needed to exist for as long as the current visit did. Notchwell's team, building later, decided a note needed to outlive the visit, since a caseload isn't a series of one-off events. I would take that first decision back the moment real, recurring caseloads showed up in the data.
I built Loomjot's note view to only ever show "now" because it was simpler to ship and the early demo looked clean with nothing cluttering the screen. It took watching a real fall-risk flag nearly slip through the cracks to understand that "clean" and "complete" aren't the same thing.
TRACE, run on two note-takersNot a design comparison. TRACE exists to find which single decision actually explains a usage pattern.
The recap, one line per letter: timeline is week five, where the curves start to split; recut is by role, case managers first; assume nothing is ruling out a transcription-quality bug; cause candidates are abandonment, workaround, and pre-editing; evidence test is the failed-lookup spike before churn.
And if you want to be sure it really works, try it somewhere elseSame five letters, a city council's meeting notes instead of a home-health visit. This time the missing history isn't personal, it's public record.
Councilmark is a civic-tech note-taker a mid-size city clerk's office uses to draft summaries of public council meetings. A different clerk drafts each week's summary, and residents sometimes ask whether an issue they raised months ago was ever addressed.
Mapped onto TRACE: the timeline shows clerk usage steady for the first two months, then a slow drop starting around month three, right when the first "has this come up before" resident question arrives. Recut shows the drop concentrated among clerks handling recurring zoning disputes, not one-off announcements. Assume nothing rules out a transcription problem, since meeting-summary accuracy stayed constant. Cause candidates are abandonment, a private spreadsheet workaround, and clerks pre-editing their own spoken summaries to avoid needing to reference anything old. The evidence test: clerks who eventually stopped using it show a spike in failed "search past meetings" attempts two to three weeks before they quietly went back to typing summaries by hand, the same signature as Loomjot's churned cohort, just on public meeting minutes instead of home visits.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "the one that keeps history wins, because nobody churns over a bad draft, they churn when they can't find last week's" and stop.
Cost: there's no budget to build full search this quarter. Say so, and start with the cheapest version: a plain list of past notes by date, no search bar, just something to scroll.
The model gets better, for real: even a Loomjot with a much more accurate transcription model still has this exact retention problem, because the missing feature was never about transcription quality.
Where people run it wrong.
They compare two products on note quality alone and miss that the real gap is what happens after the note is drafted.
They read a stable-looking average and miss that one segment is already collapsing underneath it.
They assume a quiet decline needs a quiet explanation, when a single missing feature can explain the whole thing.
How to use it live. When a retention question hands you two similar-looking products, ask yourself: what does each one let a person do three weeks after today, not just today? That question usually finds the real gap faster than comparing their launch-day feature lists.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Isn't 'add history' just an obvious feature request, not a real insight?" Response: the feature itself is obvious in hindsight, but the insight is knowing which specific moment breaks without it, which is what the failed-lookup evidence test actually proves, not just guesses at.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on AI product case study teardowns
- #1 Tear down a coding assistant: what is the core loop and where does it break?
- #2 Analyze how a major AI search product handles citation and grounding.
- #4 Tear down the onboarding of an AI product you use and identify its weakest moment.
- #5 Analyze the pricing model of an AI product and what it reveals about its cost structure.
- #6 What does an AI customer support product get right that a generic chatbot does not?
- #7 Examine an AI feature that failed publicly and identify the product decision behind it.