CaseAdvancedResponsible AI & Advanced Practice / AI product case study teardowns / #3

What product decisions explain why some AI note-takers retain users and others do not?

TRACE the products are Notchwell and Loomjot, two AI note-takers for home-health visits

Notchwell and Loomjot both listen to a home visit and draft a clinical note afterward. Colm Whitfield is a hospice social worker who carries a caseload of fourteen families and has used both, one after the other, over the same year.

The direct answer
Retention comes down to one decision: does the app let you see and search your own past notes, or does every new note quietly replace the last one? Notchwell keeps a searchable history and stayed at 70% weekly active use by week twelve. Loomjot keeps only today's draft and fell to 24% over the same stretch. Nobody churns over one bad transcription. They churn the first time they need last week's note and it's already gone.
Do this, in order
  1. Check whether the product keeps a searchable history of past notes at all.Why: this single decision is the biggest driver of the retention gap between the two.
  2. Look for when the two usage curves actually split, not just where they end up.Why: the split happened around week five, long before anyone's dashboard flagged a problem.
  3. Slice retention by role, not by an overall average.Why: case managers who need cross-visit history churned far faster than staff doing one-off intake notes.
  4. Rule out a model or transcription bug before blaming behavior.Why: transcription accuracy stayed flat the whole time; the drop was a habit change, not a quality change.
  5. Find the one action that predicts churn before it happens.Why: a failed "find my last note" search, right before someone stops opening the app, is the tell.
  6. Leave the one-off intake note flow alone.Why: a note nobody ever needs to look up again doesn't need a history feature to feel complete.

How to answer this, stage by stage

Six moves, because this question is narrower than it looks: it's really one diagnosis question wearing a comparison's clothes.

Stage 1
Scope it to two real products and one real user
Say it like this
"I'll compare Notchwell and Loomjot, two AI note-takers for home-health visits, through someone who's actually used both across the same caseload."
Why this works
Keeps "retention" from becoming an abstract debate about churn theory in general.
Stage 2
Say your structure out loud
Say it like this
"I'll use TRACE. Timeline, recut by segment, assume nothing about the cause, name three cause candidates, then find the one test that separates them."
Why this works
Signals this is a diagnosis, not a taste comparison of two apps.
Stage 3
Find when the curves actually split
Say it like this
"Both apps look almost identical for the first month. Loomjot's usage doesn't fall off a cliff, it thins out starting around week five, well before anyone would think to check a dashboard."
Why this works
A slow decline is easy to mistake for a healthy plateau if you only check the end state.
Stage 4
Recut by who actually needs the missing feature
Say it like this
"The drop isn't even. Case managers, who need to check what they wrote three visits ago, churn off Loomjot fastest. Staff doing one-time intake notes barely notice the app has no history at all."
Why this works
Shows the reversal isn't universal, it's tied to a specific real need.
Stage 5
Name the evidence test
Say it like this
"Before someone stops opening Loomjot, they try to find last week's note and fail. That failed search is the tell, and it shows up in the logs days before the person actually leaves."
Why this works
This is the concrete check, not a hunch, that separates the real cause from the other two suspects.
Stage 6
Close on the one line
Say it like this
"Nobody quits over one bad transcription. They quit the day they need their own past note and the product acts like it never existed."
Why this works
Restates the diagnosis in one breath, ready for the pushback.

Let's learn

Every evening, Colm used to spend about twenty minutes after his last visit typing up notes by hand: what he observed, what the family asked for, what he'd flagged for the nurse. It was slow, but he could always flip back through his own notebook.

Loomjot arrived first, rolled out by his agency. It listened during the visit and had a clean note waiting by the time he got to his car. For a few weeks it was simply faster. Then he needed to check something he'd noted three visits back, about a fall risk at one family's home, and Loomjot had no way to show him. Each new note replaced the last one on screen. The old one wasn't gone from a server somewhere, but there was no way to find it, search it, or even know it existed.

Knowledge spark: what's a silent decline? A drop in usage with no complaint, no support ticket, and no obvious single bad event behind it. People don't announce they've stopped trusting a tool. They just quietly stop opening it, which is exactly why it's so easy to miss until someone finally adds up the numbers.

Colm didn't complain. He started keeping his own paper tally of anything he'd want to reference later, on top of using Loomjot, which meant he was now doing two note-taking tasks instead of one.

Weekly active use, both apps, first twelve weeks
100% 50% 0 Notchwell 70% Loomjot 24% Wk 1 Wk 4 Wk 8 Wk 12 curves split here
The gap doesn't open at launch. It opens around week five, the same week case managers first need to reference a past visit.

At its worst: Colm nearly missed noting that a fall-risk flag from three visits ago had never been passed to the nurse, because Loomjot had no way to show him he'd raised it before, and he only caught it by checking his own paper tally.

The decision I would take back Loomjot's team built it to show only the current draft, with nothing kept from prior visits inside the app. That was a fine, simple choice at launch, when most early testers used it for single, unconnected appointments. It stopped making sense the moment real caseloads involved the same family across many visits, spread over months.

What I would leave alone: a one-time intake note, for a client Colm will likely never revisit, doesn't need a searchable history at all. Building that feature for every single note type would be solving a problem some notes never have.

Nobody churns over one bad transcription. They churn the day they need their own past note and the product acts like it never happened.

The lesson: a note-taker's value doesn't live in the moment it drafts a note. It lives in the moment, weeks later, someone needs to check what they said before. A product that only ever shows you today has quietly decided that moment doesn't matter.

Now here is the same thing as a story

The short version above is what you'd say defending this diagnosis live. Read this one for how the pattern actually surfaced.

Colm Whitfield has carried a hospice caseload for four years, and he can walk into a home he hasn't visited in three weeks and remember, within the first minute, exactly what he'd flagged last time.

Hand sketched timeline titled Where the two products split. Four milestones: both launch month 0, both grow month 3, Notchwell adds history month 5 highlighted, Loomjot stays flat month 9.
Nothing dramatic happens at any single point. The gap is the space between the third and fourth milestone, widening the whole time.

Loomjot's agency rollout went smoothly for the first month. Then, in week five, Colm needed to check a note from a visit three weeks back and hit a dead end: no search, no archive, no way in.

Hand sketched flow diagram titled Loomjot's one-way note pipeline. Four boxes: visit recorded, note drafted, overwritten next visit highlighted, gone for good.
Four steps, and the third one is where every prior note quietly disappears. There was never a fifth box called "find it again."

He didn't file a complaint. He started a private paper tally alongside Loomjot, the workaround of someone with too much at stake to just hope he'd remember.

Hand sketched comparison diagram titled Can he see last week's note. Left panel a document icon labeled Notchwell, caption keeps every past note. Right panel a box icon labeled Loomjot, caption shows only today's.
The two apps drafted notes about equally well. Only one of them let Colm ever see his own work again.

His agency's usage report didn't flag anything for months. Active-use percentage was still a healthy-looking number on the surface, and nobody had sliced it by role.

Hand sketched quadrant titled Sorting note-taker decisions. Axes lets you compare past runs from no to yes, and saves time up front from little to a lot. Loomjot's no history sits top left. Notchwell's run log sits top right. Paper notes sit lower middle. Voice memo sits lower left.
Loomjot sits top left: it saves real time up front and gives nothing back later. That combination is exactly what makes the decline invisible at first.

Three explanations were on the table once someone finally looked closely: staff had simply stopped using it (abandonment), staff had built their own private system around it (workaround), or staff had started keeping their spoken notes shorter and vaguer so there'd be less to forget (pre-editing).

Hand sketched icon list titled Three suspects for the silent decline. Three rows: a box icon labeled abandonment no complaint filed, a person icon labeled workaround private tally kept, a question mark box icon labeled pre-editing notes simplified first.
All three were partly true. The evidence test below is what confirmed which one actually drove people to leave.
Failed "find my last note" searches, per active user per week
4 2 0 4.1 Loomjot, later churned 0.6 Loomjot, still active 0.2 Notchwell, any user
The churned cohort hit seven times more failed lookups than users who stayed. The failed search comes before the person leaves, not after.

The evidence test settled it: the churned cohort's failed-search rate climbed for two straight weeks before their last session, while their transcription-accuracy stayed exactly as good as everyone else's. The habit didn't fade because the notes got worse. It faded because the notes became unreachable.

Loomjot's team had decided, back at launch, that a note only needed to exist for as long as the current visit did. Notchwell's team, building later, decided a note needed to outlive the visit, since a caseload isn't a series of one-off events. I would take that first decision back the moment real, recurring caseloads showed up in the data.

I built Loomjot's note view to only ever show "now" because it was simpler to ship and the early demo looked clean with nothing cluttering the screen. It took watching a real fall-risk flag nearly slip through the cracks to understand that "clean" and "complete" aren't the same thing.

TRACE, run on two note-takersNot a design comparison. TRACE exists to find which single decision actually explains a usage pattern.

T
Timeline.
The curves look identical through week four. They split starting week five, the point recurring visits begin needing a past note.
The decline started weeks before anyone's dashboard would have flagged it.
R
Recut.
Sliced by role: case managers with recurring caseloads churn fastest. Staff doing one-off intake notes barely move.
The overall average hid a specific, role-shaped drop.
A
Assume nothing.
Ruled out a model quality drop first. Transcription accuracy stayed flat the entire period, so this was a behavior change, not a bug.
A tracking or quality problem looks identical to a real usage drop on a dashboard.
C
Cause candidates.
Abandonment, workaround, and pre-editing, all three plausible, none of them generating a single support ticket.
Named three specific hypotheses instead of one vague "engagement is down."
E
Evidence test.
Failed "find my last note" searches spiked to 4.1 per week for the churned cohort in their final two weeks, versus 0.6 for users who stayed.
The single check that separates the real cause from the other two suspects.

The recap, one line per letter: timeline is week five, where the curves start to split; recut is by role, case managers first; assume nothing is ruling out a transcription-quality bug; cause candidates are abandonment, workaround, and pre-editing; evidence test is the failed-lookup spike before churn.

And if you want to be sure it really works, try it somewhere elseSame five letters, a city council's meeting notes instead of a home-health visit. This time the missing history isn't personal, it's public record.

Councilmark is a civic-tech note-taker a mid-size city clerk's office uses to draft summaries of public council meetings. A different clerk drafts each week's summary, and residents sometimes ask whether an issue they raised months ago was ever addressed.

Mapped onto TRACE: the timeline shows clerk usage steady for the first two months, then a slow drop starting around month three, right when the first "has this come up before" resident question arrives. Recut shows the drop concentrated among clerks handling recurring zoning disputes, not one-off announcements. Assume nothing rules out a transcription problem, since meeting-summary accuracy stayed constant. Cause candidates are abandonment, a private spreadsheet workaround, and clerks pre-editing their own spoken summaries to avoid needing to reference anything old. The evidence test: clerks who eventually stopped using it show a spike in failed "search past meetings" attempts two to three weeks before they quietly went back to typing summaries by hand, the same signature as Loomjot's churned cohort, just on public meeting minutes instead of home visits.

Hand sketched labeled parts diagram titled A council note-taker with a real trail. Center document icon labeled Councilmark, with four callouts: past meeting log, who edited what, compare two drafts, public comment tie.
The fix at a city clerk's office isn't different in kind from the fix for Colm's caseload. It's still: let people find what was said before.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "the one that keeps history wins, because nobody churns over a bad draft, they churn when they can't find last week's" and stop.
Cost: there's no budget to build full search this quarter. Say so, and start with the cheapest version: a plain list of past notes by date, no search bar, just something to scroll.
The model gets better, for real: even a Loomjot with a much more accurate transcription model still has this exact retention problem, because the missing feature was never about transcription quality.

Where people run it wrong.
They compare two products on note quality alone and miss that the real gap is what happens after the note is drafted.
They read a stable-looking average and miss that one segment is already collapsing underneath it.
They assume a quiet decline needs a quiet explanation, when a single missing feature can explain the whole thing.

How to use it live. When a retention question hands you two similar-looking products, ask yourself: what does each one let a person do three weeks after today, not just today? That question usually finds the real gap faster than comparing their launch-day feature lists.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits "why do some products retain users and others don't"?
Tap to flip
ANSWER
TRACE: timeline, recut, assume nothing, cause candidates, evidence test. A diagnosis question, even when it's phrased as a comparison.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Colm Whitfield, a hospice social worker with a 14-family caseload who used both Loomjot and Notchwell across the same year.
3 · THE HABIT
What did Colm start doing alongside Loomjot, once he couldn't find past notes?
Tap to flip
ANSWER
A private paper tally, kept as a workaround so he wouldn't lose track of anything he'd need to reference later.
4 · THE TENSION
What's the two-state split this whole answer turns on?
Tap to flip
ANSWER
A note-taker that keeps a searchable history, versus one where every new note quietly erases the last. Both look identical on day one.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Loomjot's decision to show only the current draft with no history, reasonable when early testers used it for one-off appointments only.
6 · THE NUMBER
Fill in the blank: by week 12, Loomjot's weekly active use had fallen to ___ percent, versus 70 percent for Notchwell.
Tap to flip
ANSWER
24 percent. Both started around 86 to 88 percent in week one.
7 · THE REPLAY
Same fall-risk lookup, a version of Loomjot with history added. What changes?
Tap to flip
ANSWER
Colm finds the three-visits-back note in seconds through search, instead of relying on a private paper tally to catch what the app couldn't show him.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which one, and what's the missing feature there?
Tap to flip
ANSWER
Councilmark, a city clerk's meeting note-taker. The missing feature is the same: a searchable log of past meetings, this time for public record instead of a personal caseload.

Check yourself Score: 0 / 0

True or false
1. True or false: Loomjot's users churned because its AI-generated notes were noticeably less accurate than Notchwell's.
  • True
  • False
Show hint
Look at the "assume nothing" step.
Show answer
False. Transcription accuracy stayed flat the whole time. The drop was about not being able to find past notes, not about note quality.
Multiple choice
2. Which single check, from the story, actually confirmed the cause of Loomjot's decline?
  • A. A survey asking users why they stopped opening the app.
  • B. A comparison of onboarding completion rates between the two apps.
  • C. A spike in failed "find my last note" searches in the two weeks before a user's final session.
  • D. A drop in the app's star rating in the app store.
Show hint
Look at the grouped bar chart of failed lookup attempts.
Show answer
C. The failed-search spike happened before churn and separated the churned cohort clearly from users who stayed.
Fill in the blank
3. Fill in the blank: churned Loomjot users averaged ___ failed lookup attempts per week, versus 0.6 for users who stayed.
Show hint
Look at the grouped bar chart.
Show answer
4.1. Roughly seven times the rate of users who kept using the app.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Showing only the current draft with no history, which made sense while early use was mostly one-off appointments with no need to look back.
Short answer, where it wouldn't matter
5. Name a kind of note where Loomjot's lack of history barely matters.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A one-time intake note for a client the worker will likely never see again. There's nothing to look back on later.
Short answer, apply it yourself
6. Pick an app you use that you've quietly stopped opening. Was there a specific thing you tried to do with it, once, that it couldn't handle?
Show hint
Think about the last time an app "failed" you quietly, with no error message, just a dead end.
Show answer
Model answer: Most people can name a specific moment, often about finding something old, not about the app being broken in the moment.
Before you close the answer
Why this works
Tests whether you can turn "why do some products retain and others don't" into a real diagnosis with a named cause and a check that proves it, instead of a list of feature differences with no evidence behind them.
Follow-up traps
"Couldn't the drop just be seasonal, fewer visits in general?" Response: no, both apps served the same caseload over the same months, and only one of the two curves dropped, which rules out a shared seasonal cause.

"Isn't 'add history' just an obvious feature request, not a real insight?" Response: the feature itself is obvious in hindsight, but the insight is knowing which specific moment breaks without it, which is what the failed-lookup evidence test actually proves, not just guesses at.
If pressed
The real fix at Notchwell doesn't just store history, it logs which specific past note a user opens during a new visit, which is what let their team confirm the "compare to a past visit" action was the one driving retention, not history in the abstract.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more