…
AI Product Case Questions

DAU of a Claude Feature Dropped 15% Last Week. Diagnose It

A worked answer to a real AI PM interview question: daily active users of a Claude feature dropped 15% last week. Diagnose it.

Transcript

Read the full transcript (1,762 words)

[INTERVIEWER] DAU of a Claude feature dropped 15% last week. Diagnose it. The daily active users of a Claude feature dropped fifteen percent last week. Diagnose it. The single worst thing you can do here is start guessing causes. Perhaps a competitor launched, or it is just seasonality, or the model got worse. That is a list of hunches, and it is an instant fail.

The right move is to isolate before you conclude. You size it, prove it is actually real, then segment until the drop points at one specific thing. A diagnosis is a search run in a fixed order, not a brainstorm of theories. This is the most important skill in the entire execution round because it is what product managers do every single week.

A number moved, and you have to figure out why in a room without the data in front of you. The interviewer wants to see discipline. Do you check the number is even real before you panic? Do you separate fewer people showing up from fewer people coming back? Do you drive all the way to the one slice that explains the delta, or do you stop at the first plausible story?

Get the order right and you will never be caught flailing at a moved metric again. Here is the method. Step one is to nail down the question before any hypothesis. Which feature exactly, and how is its daily active user metric defined? Did they just open the feature, or actually send a message in it? Is fifteen percent relative or absolute?

Was it sudden, like a cliff on one day, or gradual, like a slow bleed across the week? Is it just this feature, or is all of Claude down? Were there any known launches, holidays, or outages last week? Then comes the size check everyone skips. Is fifteen percent even outside normal variance? Plot it against the last eight weeks and against the day of the week.

If this feature normally swings ten percent week to week, then fifteen might just be noise. I would say that out loud before spending the whole interview chasing a ghost. Sizing first is what stops you solving a problem that does not exist. Step two is the reflex that scores. Assume the number is lying until it is proven honest.

Rule out the boring stuff first. Look for a logging or analytics change, where somebody redefined the event or a new app version stopped firing it. Check for a tracking bug, a timezone or holiday boundary shifting where the daily cut lands, or a data pipeline delay making the most recent days look artificially low. It could be one client, say iOS, that stopped emitting events after a release.

Here is the concrete cross check. Pull a second independent source. Do the server side request logs for the feature agree with the product analytics daily active user count? If server logs are flat but the dashboard dropped, this is a measurement break, not a user exodus. The entire diagnosis changes direction right there. Never skip this. More daily active user down alerts are broken counters than real crises.

Step three. If the drop is real, split the cause space in half so you are not searching everywhere at once. Internal means something we did. That could be a release last week, a config or paywall change, a pricing move, an entry point that got moved or broken, a model swap that hurt quality, or a latency regression that made the feature slow and annoying.

External means something in the world. Think seasonality, a holiday week, summer, a competitor launch, an upstream dependency changing, an API or partner surface that used to drive traffic, or a platform change like an operating system update or an app store issue. I would check internal first because internal causes are both more likely and crucially checkable. You have the release timeline right there.

Correlate the start date of the drop against what shipped. That alignment is often the whole answer. Step four is the core of the whole thing. Break daily active users down along every axis and find where the delta concentrates. Here is the key insight. A uniform fifteen percent across everything means a broad cause like measurement, pricing, or seasonality.

A drop concentrated in one slice hands you the cause on a plate. So I go through the axes. Platform means iOS versus Android versus web. If iOS alone dropped forty percent and everything else is flat, I am looking at the last iOS release. App version asks if the drop started exactly when version X rolled out. That is close to a smoking gun.

Geography and language means one region only, which points at a local outage, a regional launch, or a regional holiday. Then cohort, and this one is underused. Decompose the metric into new plus retained plus resurrected. Did we stop acquiring new users due to a broken signup or a paused marketing campaign, or did retained users leave due to a quality or experience regression?

Those have completely different fixes, and lumping them together as a general drop hides which one broke. Finally, user segment covers free versus paid, new versus tenured, and heavy versus light. A paid only drop right after a change smells like a pricing or entitlement bug. The goal through all of this is to keep slicing until one segment explains most of the delta.

Step five. Within that drop, split the funnel. Are fewer users arriving at the feature, indicating an entry point, discovery, notification, or acquisition problem? Or are the same users arriving but fewer engaging once they are there, meaning the feature got slower, worse, or broke? Walk the actual steps from impression of the entry point, to open, to first action, to success.

Find the exact step where the curve cracks. Fewer arriving points you upstream to a moved button, a marketing change, or a broken link. Fewer engaging points you at the feature itself, highlighting latency, quality, or a bug in the flow. That single split of arriving versus engaging cuts your remaining hypotheses roughly in half. Step six is to close it out.

Pull the single segment that explains most of the delta and correlate it against the release timeline and the latency and error dashboards to confirm cause, not just location. Location tells you where, cause tells you why, and you need both before you act. Then act by cause. If it is a regression from a release, roll it back. For a measurement artifact, fix the logging and restate the number.

For an external or seasonal cause, confirm it against prior years and adjust the forecast rather than firefighting something you cannot control. Then communicate the finding, your confidence level, and the action. Only now do you actually have a diagnosis because you isolated before you concluded. Let me run the whole method on a real case. The feature is Claude Projects.

Daily active users, defined as users who open a project, are down fifteen percent relative week over week. Size it. The last eight weeks swing about five percent, so fifteen is real signal, not noise, and it is a step down starting last Wednesday, not a gradual bleed. Real or artifact. The server side request logs for the feature also show the drop, so it is real, not a broken event.

Internal versus external. A new web release shipped last Wednesday, the exact drop date, which is suspicious. Segment. The drop is almost entirely on web, down thirty five percent, while iOS and Android are flat. Cohort. It is retained users, not new, so acquisition is fine and existing users just stopped coming back. Funnel. Those same users are still arriving at the project entry point and impressions are flat, but open to first action collapsed, so they arrive and immediately bounce.

Confirm. Checking the latency dashboard, p95 load time for Projects on web jumped from 1.2 seconds to six seconds starting Wednesday, and the error dashboard shows a spike in a specific API call that the new release introduced. Cause found. The web release added a slow, sometimes failing call on project open. Act. Roll back that change, latency returns to 1.2 seconds, and daily active users recover within two days.

Here is the contrast that proves the method. If the server logs had been flat instead of down, the answer would have been a measurement break in the web event pipeline, meaning we fix the logging and restate the number. That is a completely different conclusion and a completely different team fixing it. That is exactly why you isolate before you blame.

Here is what makes them lean in. First, you isolated before concluding. You sized the drop, proved it was real against a second data source, and segmented until it pointed at one change instead of reciting a list of possible causes. Second, you decomposed the metric into new, retained, and resurrected, and read the funnel for arriving versus engaging, which separates acquisition breaking from the experience breaking.

Third, you correlated the isolated segment against the release timeline and the latency and error dashboards to confirm the cause, not just the location. That discipline is the single strongest signal in this round. Now for the traps. Trap one is jumping straight to hypotheses like maybe a competitor launched before you have even checked the number is real or outside normal variance.

Trap two is treating a daily active user drop as identical to a retention drop when it could just as easily be an acquisition drop, a resurrection drop, or a pure logging artifact. Those are three different problems with three different owners. Trap three is segmenting on one axis, say geography, finding nothing obvious, and stopping instead of driving all the way to the single slice that explains most of the delta.

Stopping early is how you walk out with a guess that it is probably seasonal and no actual diagnosis. So the whole method in order is to clarify and size it, prove it is real against a second source, split internal versus external, segment by platform, version, geo, and cohort until one slice concentrates the delta, read the funnel for arriving versus engaging, then confirm cause against the release and latency dashboards and act.

Carry this into the room. Diagnosis is a search run in order. Size it, prove it is real, then segment by platform, version, geo, and cohort until the drop points at one change. Find the deploy.

Keep learning