Artifact critiqueAdvancedAI Opportunity & Model Strategy / Data strategy as product strategy / #20

Write the section of a strategy doc that argues for a data investment with no immediate feature payoff.

PICK the sprint that doesn't ship a feature, argued in the doc itself

Picture the strategy doc before anyone has argued for this line item. Wren Relief Network runs AskWren, a text line that matches people in crisis with nearby food banks and shelters. Marisol Adeyemi is Head of Product, and the doc she's writing has to convince a skeptical board that the next sprint should go to a data pipeline nobody will ever see, instead of the Spanish-language rollout everyone is already asking for.

The direct answer
Argue for the investment by naming which cost is hidden and which one is visible, then pick the hidden one on purpose. Say plainly: delaying the visible feature costs three weeks that everyone will notice and forgive. Skipping the outcome-verification pipeline costs a truth we can never go back and recover, because every month without it is a month of referrals we can never re-check. Set a real kill criterion so the investment doesn't run forever on faith.
Do this, in order
  1. State your position before your reasoning: invest in verification first, delay the feature.Why: a doc that hedges until paragraph four reads as uncertain, not careful.
  2. Name who feels each cost, in real terms, not abstractions.Why: "a delay" and "a data gap" mean nothing until you say who notices each one and when.
  3. Say which cost is hidden and unrecoverable, and commit to optimizing against that one.Why: the visible cost gets fixed by itself once you ship. The hidden one only gets worse the longer it's ignored.
  4. Write a kill criterion into the doc itself.Why: without one, "invest in data" becomes a standing tax nobody ever revisits.
  5. Say plainly when the feature really should win instead.Why: shows judgment, not a reflex to always pick the unglamorous option.

How to answer this, stage by stage

Nobody is scoring your writing style. They're scoring whether you can make an invisible cost feel as real as a missed launch date.

Stage 1
Scope it to one real ask
Say it like this
"I'll ground this in AskWren at Wren Relief Network, a text line matching people in crisis to nearby aid, and the actual sprint fight over Spanish-language support versus an outcome-verification pipeline."
Why this works
Keeps the doc from turning into a general essay about "the importance of data."
Stage 2
State your position before your reasoning
Say it like this
"I'll run this as PICK. Position first, then impact, then the cost asymmetry, then a kill criterion so this isn't a bet with no exit."
Why this works
Signals you can commit to a call, which is exactly what a skeptical board is testing for.
Stage 3
Reframe: this isn't data versus features, it's visible versus hidden
Say it like this
"This isn't really 'ship a feature or build a data project.' It's 'accept a cost everyone can see, or accept one that's invisible until it's a news story.' Framed that way, the choice gets a lot clearer."
Why this works
This is where a strong answer separates from someone who just says "data is important."
Stage 4
Give the actual line that goes in the doc
Say it like this
"Here's the sentence I'd put in bold: 'We currently cannot tell the difference between a referral that helped someone and a referral that sent them to a closed building. Every week we ship features on top of that gap, we make it more expensive to fix.'"
Why this works
This is the direct answer, written the way it would actually appear on the page, not summarized.
Stage 5
Prove it with what already slipped through
Say it like this
"Last quarter's sample review found 22 percent of referrals marked 'successful' actually pointed to a shelter or food bank that was full or closed. Our dashboard read 98 percent matched the whole time. Nobody caught it until a local reporter did."
Why this works
Compresses the whole argument into the one gap between what the dashboard said and what was actually true.
Stage 6
Close with the kill criterion, and the AI-specific reasoning
Say it like this
"The honest reason this isn't just 'more testing' is that our matching model is only ever as good as the labels we feed back into it, and right now our only label is 'message sent,' not 'help received.' I'd commit to this: once sampled success holds above 90 percent for a full quarter, we stop investing here and ship the backlog instead."
Why this works
Closes with a real exit condition, not an open-ended promise to "keep investing in data."

Let's learn

Here is what happens when the number on your dashboard and the truth on the ground quietly stop being the same thing.

AskWren launched two years ago handling about 120 texts a day: someone in crisis writes in, the model matches them to a nearby shelter or food bank, and a reply goes out with directions and a phone number. Today it handles about 900 texts a day. Back at 120, Renata Sokol, the crisis-line coordinator, personally called back a sample of harder cases within 48 hours, just to see if the resource had actually helped. At 900 a day, she can't. Somewhere around month fourteen, she quietly stopped.

Hand sketched labeled parts diagram titled What the callback pipeline actually needs. A document icon at the center labeled Outcome Pipeline, with four labeled callouts: 48 hour auto text, Partner capacity ping, Human sample review, Owner Renata.
None of these four parts exist yet. That's the whole gap the doc is arguing to close.

Here's the turn: "message matched" kept reading 98 percent the entire time. Nobody had a number for "message that actually helped," because nobody had ever built the pipeline that would tell them the difference.

What the dashboard says versus what a hand sample found
100% 50% 0 98% Dashboard: matched 78% Hand sample: actually helped
The 20-point gap between these two bars is the entire argument for this sprint.

At its worst, an organization can look flawless on paper for years while a fifth of the people it's supposed to help are quietly being sent to a door that won't open, and the only thing standing between "looks fine" and "front-page story" is which week a reporter happens to call.

The choice I would take back Early on, when a leader told the board "every roadmap item ships something users can see," that promise made sense while AskWren was small enough for Renata to personally check the hard cases by hand. It stopped making sense once volume outgrew what one person could quietly cover for free.

What I would leave alone: I wouldn't delay the Spanish-language rollout by more than a sprint. That's a real, visible cost to real people waiting for it, and PICK only wins here because the hidden cost is worse, not because visible costs never matter.

The lesson: a metric that only counts what your system did, never what happened to the person after, will always look better than it deserves to. The gap doesn't show up on a dashboard. It shows up in a news story, months late.

Now here is the same thing as a story

The short version above is what you'd write into the doc itself. Read this one for how a sensible promise quietly turned into a blind spot over a year.

Her name is Renata Sokol. She has worked crisis intake for six years, four of them at Wren. Ask her what a caller needs and she'll usually know before they finish the first sentence.

When AskWren launched, Renata treated it as a second pair of hands, not a replacement for her judgment. Every evening, she'd pull ten of the day's harder matches and call the person back: did the shelter have a bed? Was the food bank actually open? For the first year, at 120 texts a day, that was easy to keep up with.

Strategy doc excerpt, data investment section
To: Wren Relief Network leadership · Re: Q3 roadmap, outcome-verification pipeline · From: Marisol Adeyemi, Head of Product

We are proposing to spend the next sprint building a lightweight callback and partner-capacity pipeline instead of shipping Spanish-language support, which we know is wanted and ready to build.

Here is the honest reason. AskWren currently measures success as "a matched resource was texted back." It does not measure whether that resource was open, had room, or actually helped. A hand sample of last month's referrals found that 22 percent pointed to a shelter or food bank that was at capacity or closed at the time. Our own dashboard read 98 percent matched the entire month.

This gap does not show up as a feature request. It shows up, eventually, as a story about someone we failed while our own numbers said we were succeeding. We can close it now for the cost of one sprint, or we can keep shipping features on top of a foundation we cannot currently trust.

We are not asking for open-ended investment. Once sampled referral success holds above 90 percent for a full quarter, we recommend returning fully to the feature roadmap.

For most of that year, the good months were genuinely good. Then AskWren's reach grew, word spread, and daily volume climbed past 900. Renata kept trying to call back her ten hardest cases every night. Then it was five. Then, without ever deciding to, she stopped.

Hand sketched comparison titled The asymmetry, drawn. Left panel, a document icon labeled Ship Spanish support now, caption three weeks sooner, a visible win. Right panel, a question mark box icon labeled Skip the verification pipeline, caption nobody sees the gap until it's a headline, shown in a different color.
One box is small and easy to defend in a meeting. The other one is the whole reason this doc exists.

Nobody at Wren decided, on any single day, that referral quality no longer needed a human check. It just quietly stopped being anyone's job once the one person who used to do it for free ran out of hours in a day.

Knowledge spark: what's a proxy metric, and why does it lie? A proxy metric stands in for the thing you actually care about because the real thing is harder to measure. "Message sent with a match" is a proxy for "person got real help." Proxies are fine right up until the gap between them widens quietly, with nothing on the dashboard telling you it's happening.

The gap widened for eleven months before a local reporter, working on a piece about a shelter that had closed two months earlier, found that AskWren was still texting people its address. The story ran the same week the board was reviewing next quarter's roadmap.

Hand sketched metaphor scene titled One cost you can see, one you can't. Left, a scale icon labeled THREE WEEKS, caption a delay everyone can see and forgive. Right, a funnel icon labeled NO GROUND TRUTH, caption a gap nobody can go back and fill, shown in a different color.
A three-week delay is a scale that balances back out. A missing quarter of ground truth never does.
A delayed feature is a cost with an end date. A silent gap between what your dashboard says and what actually happened has no end date at all, only a day someone finally notices it.

Two years earlier, when a leader first told the board "every roadmap item ships something a user can see," it sounded like the right kind of discipline for a young nonprofit trying to prove it could deliver. Nobody meant it to become a rule against ever funding something invisible.

Hand sketched icon list titled What the doc section actually argues. Three rows: a document icon, message sent is not the same thing as help received. A funnel icon, every month with no outcome data is one we can never reconstruct. A gauge icon, a three week delay is visible, a silent truth gap is not.
Three sentences, and the whole doc section is really just these, said plainly.
Sampled referral success, after the pipeline ships, against the kill line
100% 50% 0 Kill line: 90% Before: 78% Q1: 84% Q2: 90% Q3: 93%
By quarter two, sampled success crosses the kill line, exactly the signal that says stop investing here and return to the feature roadmap.

Rerun the same year with the pipeline built from the start: a 48-hour automated text checks in on every referral, a lightweight partner API confirms real-time capacity, and Renata reviews a flagged sample instead of trying to cover all of it herself. The reporter's story never runs, because the closed shelter gets flagged and pulled from matching within two days of closing, not two months.

What I'd tell myself, reading that reporter's first email to our press line: the promise to always ship something visible was never wrong to make. It was wrong to leave standing long after the org had outgrown the one person quietly making it true for free.

PICK, the sentence that goes in the docNot a lecture on why data matters. PICK is what turns "we should invest in data" into an actual, defensible line item.

P
Position. Your pick, before any reasoning.
Invest in the outcome-verification pipeline this sprint. Delay Spanish-language support by three weeks.
This is the hardest sentence to say plainly in a room that wants a shiny roadmap, and it's the one the whole doc rests on.
I
Impact. Who feels each kind of cost.
Spanish-speaking users feel three weeks of waiting. People sent to closed or full resources feel a much worse cost, one nobody currently measures at all.
Naming both sides in real terms is what keeps this from sounding like an abstract values statement.
C
Cost asymmetry. One is cheap and visible, the other hidden and compounding.
A delayed launch is a cost with an end date. A month with no outcome data is a month of referrals we can never go back and check.
This is the argument. Optimize against the cost that keeps growing quietly, not the one everyone can already see.
K
Kill criteria. What would flip this pick.
Sampled referral success holding above 90 percent for a full quarter means stop investing here and return fully to the feature roadmap.
Without this line, "invest in data" has no exit, and starts to sound like faith instead of a decision.

The recap, one line per letter: position is invest in verification, delay the feature by three weeks; impact is Spanish-speaking users waiting versus people sent to closed doors with nobody measuring it; cost asymmetry is a bounded delay against an unrecoverable truth gap; and kill criteria is 90 percent sampled success held for a quarter, which ends the investment on purpose instead of letting it run forever.

And if you want to be sure it really works, try it somewhere elseSame four letters, a funeral home's scheduling tool instead of a crisis line. Different flip family entirely, the same quiet promise breaking.

Thornwell Funeral Home uses a small AI tool that drafts service timelines and staffing plans from a family's intake form. Mapped onto PICK: position is invest a few weeks in tracking whether families' actual requests changed mid-process, instead of shipping a requested online payment feature next. Impact is families waiting slightly longer for online payment, against grieving families whose last-minute changes never make it back into the model because nobody logs them. Cost asymmetry is a payment feature delay that ends the day it ships, against silently repeating the same scheduling mistakes on every family whose needs shift, because the system never learns what changed and why. The flip here is delegation, not abandonment: Wendell Ashgrove, the senior director, used to let junior staff run the intake tool unsupervised once it proved reliable. When it started missing late changes to services, he started re-reviewing every single intake himself, and two people were now doing the work one tool was supposed to do alone.

Hand sketched decision tree titled When the investment has paid for itself. Root: Do we actually know if referrals help. Three branches: sampled success under 85 percent leads to keep investing, sampled success 85 to 90 percent leads to automate and reassess, sampled success over 90 percent stable leads to stop, ship features.
The same three branches work for a funeral home's intake changes as they do for a crisis-line referral.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "pick the cost that's hidden and can't be undone over the one that's visible and ends on its own," and stop.
Cost: no sprint to spare for either. Say so honestly, and ship the smallest possible version, an automated 48-hour text with no human review yet, rather than nothing at all.
The feature turns out cheap to delay for real: if the visible feature can slip two weeks with zero real-world cost, that's a legitimate reason to invest in the hidden gap first, not a shortcut you're taking.

Where people run it wrong.
They treat "ship something visible every sprint" as an unbreakable rule instead of a promise that made sense at one size and not another.
They let a proxy metric stand in for the real outcome forever, instead of setting a date to go check whether the proxy is still telling the truth.
They invest in a data pipeline with no kill criterion, turning a smart bet into a standing tax nobody ever revisits.

How to use it live. The moment someone asks you to defend an invisible investment, ask yourself: which of these two costs can I actually undo later, and which one compounds quietly while I wait? Name the one that compounds, and the doc practically writes itself.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Abandonment flip: Renata quietly stopped personally calling back hard cases once volume outgrew what she could cover, with no complaint or ticket marking the moment it happened.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Renata Sokol, a six-year crisis intake worker who used to personally verify AskWren's harder referrals by phone.
3 · THE HABIT
What did Renata stop doing because the tool seemed to be working?
Tap to flip
ANSWER
She stopped calling back the day's harder referrals to check whether the matched resource actually had room or was even open.
4 · THE ASYMMETRY, IN THIS STORY
What are the two costs being weighed against each other?
Tap to flip
ANSWER
A three-week delay to Spanish-language support, visible and bounded, against a silent gap between "matched" and "actually helped" that grows worse the longer it's left unmeasured.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Telling the board "every roadmap item ships something visible," a promise that made sense while the org was small enough for one person to cover the gap for free.
6 · THE NUMBER
Fill in the blank: a hand sample found ___ percent of "successfully matched" referrals actually pointed to a closed or full resource.
Tap to flip
ANSWER
22 percent, against a dashboard reading 98 percent matched.
7 · THE REPLAY
Same year, the outcome pipeline built from the start. What changes?
Tap to flip
ANSWER
A closed shelter gets flagged and pulled from matching within two days instead of two months. Sampled success climbs past the 90 percent kill line by quarter two, and the reporter's story never runs.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Thornwell Funeral Home's intake tool. The flip is delegation: Wendell Ashgrove reclaimed every intake review himself once the tool started missing late changes, doubling the work it was meant to remove.

Check yourself Score: 0 / 0

True or false
1. True or false: the doc excerpt argues that Wren should never ship a visible feature again until every data problem is fixed.
  • True
  • False
Show hint
Look at the kill criteria step and "what I would leave alone."
Show answer
False. The doc sets a specific kill criterion: once sampled success holds above 90 percent for a quarter, investment stops and features resume.
Multiple choice
2. Why does this answer choose to delay the Spanish-language feature instead of the verification pipeline?
  • A. Spanish-language support is less important to the mission.
  • B. The feature's cost is visible and bounded; the data gap's cost is hidden and compounds the longer it's ignored.
  • C. Engineering estimated the pipeline would take less time to build.
  • D. The board specifically requested the pipeline first.
Show hint
Look at the cost asymmetry step and the metaphor scene diagram.
Show answer
B. The whole argument turns on optimizing against the cost that keeps quietly getting worse, not the one that ends the day you ship.
Fill in the blank
3. Fill in the blank: the dashboard read 98 percent matched, but a hand sample found the real "actually helped" rate was closer to ___ percent.
Show hint
Look at the bar chart comparing the dashboard number to the hand sample.
Show answer
78 percent. A 20-point gap between what the dashboard claimed and what a manual check actually found.
Short answer, where it wouldn't matter
4. Name a situation where delaying a visible feature for a data investment would NOT be the right call, according to this answer's own logic.
Show hint
Look at "what I would leave alone" and the kill criteria step.
Show answer
Model answer: If sampled referral success were already holding above 90 percent for a full quarter, there'd be no hidden gap left to justify delaying the feature further.
Short answer, apply it yourself
5. Think of a product you use that reports a success rate. What real outcome might that number be standing in for, and how would you know if the two had quietly split apart?
Show hint
Look for the difference between "the system did its part" and "the person actually got what they needed."
Show answer
Model answer: A job board's "application submitted" rate stands in for "person got hired." You'd only catch the gap by following up with a sample of applicants directly, since the submission count alone can't tell you.
Short answer, work the number
6. If Wren handles 900 texts a day and 22 percent of "matched" referrals are actually bad, roughly how many people a day were being sent somewhere that couldn't help them?
Show hint
22 percent of 900.
Show answer
Model answer: About 198 people a day, roughly 1,400 a week, sent to a resource that was full or closed while the dashboard still read 98 percent matched.
Before you close the answer
Why this works
Tests whether you can argue for an invisible investment in concrete, bounded terms, or whether you fall back on "data is important" without ever naming the actual cost of not having it.
Follow-up traps
"Couldn't you just ship the feature and add verification later?" Response: every week without it is a week of referrals we can never go back and re-check, since the people involved have already moved on by the time we'd look.

"What if the sample size for that 22 percent number is too small to trust?" Response: fair, which is exactly why the doc proposes building a real pipeline instead of running on hand samples forever.
If pressed
The pipeline that shipped used a partner API ping for real-time capacity plus a templated 48-hour follow-up text, and it cut Renata's manual review load to a flagged 15 percent of referrals instead of zero or all of them.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more