CaseIntermediateAI Opportunity & Model Strategy / Data strategy as product strategy / #4

Your company has ten years of unstructured documents. Is that an asset? Interrogate the claim.

PICK · the audit that found Averlynn Wealth Partners' notes had gone quiet

Averlynn Wealth Partners is an independent advisory firm with about 240,000 client meeting notes going back ten years. Everett Sowande leads data and compliance there. When a proposal came in to fine-tune a "next-best-action" copilot on that archive, he ran an audit before anyone touched a model.

The direct answer
Ten years of documents is not an asset by itself. It is a claim, and you have to test it against three things: whether you actually have the right to use it this way, whether it is honest enough to teach anything true, and whether it is structured enough to use at all. Until those three checks pass, call the archive raw material, not an asset, because assuming otherwise is the expensive mistake here, not the cautious one.
Do this, in order
  1. Test the actual right to use the archive this way, before anything else.Why: a consent or recordkeeping violation is the one mistake here you cannot undo once a model has trained on it.
  2. Sample-audit the archive for honesty and specificity, not just for volume.Why: notes written to protect the writer, not to record the truth, teach a model the wrong lesson with total confidence.
  3. Check whether the archive is actually structured, or just years of free text.Why: raw prose without consistent fields is expensive to use no matter how large it is.
  4. Treat the volume, ten years, hundreds of thousands of records, as irrelevant until the first three checks pass.Why: a big number of unusable records is still unusable.
  5. Only then, decide whether to clean the archive or start fresh with a smaller, better-instrumented dataset.Why: the honest comparison is clean-and-small against dirty-and-large, not "we already have it, so let's use it."
  6. Revisit the archive's real value whenever a legal or regulatory event might have changed how honestly it was written.Why: what the notes actually contain can quietly change without anyone touching the file format.

How to answer this, stage by stage

Nobody is scoring whether you know that data can be valuable. They're scoring whether you can name the exact three checks that separate a real asset from a large pile of risk.

Stage 1
Scope the claim to one archive
Say it like this
"Let's ground this in a real archive, Averlynn Wealth Partners' ten years of client meeting notes, and the exact audit that got run before anyone tried to fine-tune anything on it."
Why this works
Keeps the answer from becoming an abstract debate about whether "data is valuable" in general.
Stage 2
Say your structure out loud
Say it like this
"I'll run this as PICK. Position, my actual call, stated first. Impact, who feels each kind of error. Cost asymmetry, which error is hidden and expensive. Kill criteria, what evidence would change my mind."
Why this works
Signals you're committing to a position instead of hedging with "it depends," which is exactly what this question is testing.
Stage 3
Reframe: the real question is what it costs to be wrong
Say it like this
"This isn't really a question of how much data we have. It's a question of what it costs us if we're wrong about it being ready to use, versus what it costs us to double-check first."
Why this works
This is where a strong answer stops sounding like enthusiasm about "big data" and starts sounding like judgment.
Stage 4
Give the position
Say it like this
"My position: it's raw material, not an asset, until we've checked rights, honesty, and structure. Ten years of volume proves nothing about any of those three."
Why this works
This is the sentence an interviewer should be able to quote back at you five minutes later.
Stage 5
Prove it with the compressed failure
Say it like this
"At Averlynn, only 14 percent of ten years of notes carried any usable structure by the time anyone checked. And the notes themselves got quietly vaguer after year six, once a detailed note became evidence in a client dispute. Average note length dropped from 340 words to 61 almost overnight."
Why this works
This is the number the whole "we already have ten years of data" claim would fall apart without.
Stage 6
Name the kill criteria
Say it like this
"If a compliance review confirms the notes are cleared for this use, and a sample audit shows consistent structure across a real chunk of the archive, I'd change my answer. Right now, neither is true."
Why this works
Shows this is a real, testable position, not a fixed opinion nothing could ever move.
Stage 7
Close on one line
Say it like this
"Ten years of documents isn't an asset. It's ten years of a claim nobody had tested until someone finally opened the file and counted."
Why this works
Restates the direct answer, letting the number do the closing work.

Let's learn

Here is what happens when nobody checks whether an archive that sounds valuable is actually honest.

Averlynn Wealth Partners keeps a note after every client meeting: what was discussed, what was recommended, how the client responded. Ten years of that adds up to about 240,000 records. In the first three years, the CRM had a required dropdown, recommendation accepted, declined, or deferred, and 89 percent of notes carried that structured tag alongside the free text.

Hand sketched icon list titled PICK in one screen. Four rows: Position, your pick, before any reasoning, scale icon. Impact, who feels each kind of error, person icon. Cost asymmetry, which error is hidden and expensive, gauge icon, in a different color. Kill criteria, what evidence flips the pick, question mark box icon.
The four letters, held up as one page. Cost asymmetry is the step this question is really testing.

A redesign in year three removed that dropdown to cut clicks from the meeting-close flow, leaving only the open text box. By year ten, only 14 percent of notes still carried anything resembling a structured, checkable outcome.

Here's the turn: the missing dropdown wasn't the real story. The real story is what happened to the free text itself. In year six, a detailed note became central evidence in a client dispute, quoted line by line in a deposition. After that, notes across the firm got shorter, vaguer, and far more careful, without anyone deciding this on purpose.

Share of sampled notes with a usable structured outcome, by year
100% 50% 0 40% bar Yr 1 Yr 7 Yr 10 14%
The line crossed below a reasonable "ready to use" bar around year four, three years before the lawsuit even happened.

At its worst, a model trained on this archive learns two lessons nobody wanted to teach it: which advisors write the vaguest, most defensive notes, and how to sound reassuring without actually recommending anything specific, since that's exactly the pattern the post-lawsuit notes reward.

Ten years of documents isn't an asset. It's ten years of a claim nobody had tested, sitting quietly in a shared drive until someone finally opened the file and counted.
The choice I would take back Removing the required "recommendation outcome" dropdown in year three, to cut a click from a busy advisor's day. That made sense when the goal was a faster meeting-close flow. It stopped making sense the moment the only remaining record of ten years of client advice was whatever an advisor felt like typing that day.

What I would leave alone: I wouldn't force detailed structure onto informal internal scheduling notes on the same CRM, since nobody was ever going to train a client-facing model on "moved to Thursday, client running late."

The lesson: an archive's age and size tell you nothing about whether it's honest. Only a sample audit tells you that, and the audit is worth running before the roadmap, not after the model.

Now here is the same thing as a story

The short version above is what you'd say defending a "don't fine-tune on this yet" recommendation in a steering committee. Read this one for how a required field's removal quietly changed what ten years of notes actually contain.

Everett has led data and compliance at Averlynn for six years. He wasn't the one who removed the dropdown. He was the one who had to explain, years later, why it mattered that someone had.

Hand sketched timeline titled The ten years, before and after the lawsuit, third milestone emphasized. Four milestones: Structured field removed, year 3, to reduce clicks. Notes drift to free text, years 4 to 6. A note becomes evidence, year 6, a client dispute, shown in a different color. Notes go vague firmwide, years 7 to 10.
Two separate decisions, three years apart, compounded into one archive nobody could fully trust by year ten.

The proposal that landed on his desk was simple on paper: fine-tune a copilot on ten years of meeting notes so it could suggest a next best action for any client, using the firm's own accumulated judgment. The pitch deck's first slide said "240,000 records of proprietary advisor expertise." Nobody in that room had actually read a random sample yet.

Hand sketched comparison titled The asymmetry, drawn. Left panel, a document icon labeled SLOWER START, caption clean data, smaller, a visible cost. Right panel, a question mark box icon labeled HIDDEN VIOLATION, caption PII and liability buried in the notes, shown in a different color.
One box is small and honest about its cost. The other is the one that quietly ends careers if it's wrong.

Everett pulled two hundred notes at random, evenly spread across all ten years. The pattern by year was obvious within an afternoon. Years one through three read like real clinical judgment: "Client expressed concern about market volatility; recommended shifting 10% from equities to bonds; client accepted." Years seven through ten mostly read like: "Met with client. Reviewed portfolio. No changes needed at this time."

Knowledge spark: why would a lawsuit make notes worse, not better? Once a detailed note gets read back in a deposition, the person writing the next note isn't thinking about training a future model. They're thinking about the next lawyer who might read it. Specificity starts to look like exposure. Vagueness starts to look like safety. The record gets less useful at exactly the moment everyone is being more careful.
Hand sketched metaphor scene titled What the note actually protects. Left, a person icon labeled THE ADVISOR, caption a vague note protects them, not the client. Right, a document icon labeled THE ARCHIVE, caption loses the exact detail a model needs, shown in a different color.
Both things are true at once. A vaguer note can be the right call for the advisor and the wrong outcome for the archive.

Everett found the year-six lawsuit in an old compliance memo, not in the CRM itself. The note quoted in the deposition had specified an exact percentage shift and an exact reason, and a client later argued that reason had been wrong. Nobody blamed the note-taking. Everyone quietly changed how they wrote notes afterward anyway.

Hand sketched labeled parts diagram titled What a usable meeting record needs. A document icon at the center labeled Usable Note, with four labeled callouts around it: Recommendation given. Client response. Risk tolerance noted. Follow up outcome.
Four parts. Years one through three had all four most of the time. Years seven through ten had almost none of them.

When the dropdown was removed in year three, someone said, "let's simplify the meeting-close screen, advisors are already writing it in the notes anyway," and it sounded completely reasonable, since at the time, most of them genuinely were.

Estimated cost of each path, if the assumption is wrong
$2.5M $1.25M 0 Raw archive, if wrong $2.4M Safer rebuild path $180K
One bar is a known, budgeted delay. The other only shows up on the balance sheet if you find out the hard way.

Rerun the same ten years with the dropdown kept mandatory, and a firm policy that a note's specificity is protected, not penalized, in any legal review: the archive still shows the same year-six dispute, but the notes around it stay just as detailed as the years before, because nobody ever learned that detail was a liability. A sample audit ten years later finds close to 80 percent of records usable instead of 14, and the copilot project starts from an actual asset instead of a pile of good intentions.

What I'd tell myself, reading two hundred notes that got steadily vaguer for no reason anyone had chosen on purpose: the archive didn't fail because advisors got worse at their jobs. It failed because nobody protected the one thing that made it valuable, which was advisors feeling safe enough to write down exactly what they actually did.

PICK, the call that would have caught this before the pitch deckNot a definition of what makes data valuable. PICK is what forces you to commit to a position and name what would change it.

P
Position. Your pick, before any reasoning.
Raw material, not an asset, until rights, honesty, and structure are all checked.
Commits to a real answer instead of "it depends," which is what this exact question is built to test.
I
Impact. Who feels each kind of error.
Assuming the archive is ready hits clients whose private details get baked into a model without real consent. Assuming it isn't ready just costs the team a slower start.
Names both sides honestly instead of only describing the cost of caution.
C
Cost asymmetry. Which error is hidden and expensive.
Training on an unvetted archive risks a hidden regulatory or liability exposure, roughly $2.4M in this estimate, against a visible $180K delay from rebuilding clean.
This is the hardest step, and the one a pitch deck slide with a big record count always skips.
K
Kill criteria. What evidence flips the pick.
A confirmed compliance clearance plus a sample audit showing real structure across a meaningful slice of the archive.
Makes this a testable claim, not a permanent verdict against the whole archive.

The recap, one line per letter: position is treating the archive as raw material until proven otherwise, impact is naming both the client's hidden exposure and the team's visible delay, cost asymmetry is the roughly $2.4M hidden risk against a $180K known cost, and kill criteria is a compliance clearance plus a real structure audit, either of which would flip the call.

And if you want to be sure it really works, try it somewhere elseSame four letters, a veterinary practice instead of a wealth manager. Different flip family entirely, the same interrogated claim.

Dr. Wendeline Ashcroft runs Copperlatch Veterinary, which has ten years of handwritten treatment charts, now scanned, that a vendor wants to fine-tune into a treatment-recommendation assistant. Mapped onto PICK: position is that the charts are raw material until their consistency is proven, not an asset by page count. Impact is a bad recommendation reaching an animal whose owner has no way to catch a subtle dosing error, against the slower, visible cost of building from two years of clean digital records instead. Cost asymmetry favors caution, since a wrong dosing suggestion is the hidden, expensive error. Kill criteria is a sample audit confirming the charts consistently record the actual clinical reasoning, not just a diagnosis code.

The flip here is delegation, not concealment. As Copperlatch grew, senior vets started handing chart-writing to vet techs to save time during busy shifts, then quietly took detailed charting back for themselves once they noticed tech-written charts recorded the diagnosis and treatment but never the reasoning behind choosing one drug over another, which is exactly what a future recommendation model would need to learn from.

Hand sketched quadrant titled Sorting ten years of charts by what's usable. X axis specificity, vague to specific. Y axis consistency, erratic to consistent. Early vet written charts placed specific and consistent. Recent tech written charts placed vague and erratic. Post scare charts placed vague but consistent.
Three very different clusters hiding inside one ten-year archive, none of them visible from the total page count alone.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "it's raw material until rights, honesty, and structure are checked, not an asset by default," and stop.
Cost: no budget for a full compliance review before a pilot needs to start. Say so honestly, and run the cheap sample audit first, since that alone often settles the question.
The archive turns out to be genuinely fine, for real: if the sample audit shows strong, consistent, cleared records, that's a legitimate reason to proceed with confidence, not a shortcut being taken.

Where people run it wrong.
They treat record count as a proxy for data quality, without ever sampling the actual content.
They assume old records were captured under consent that covers a brand new use, like model training, without checking.
They discover the archive's real condition only after a model has already trained on it, instead of auditing first.

How to use it live. The moment an interviewer hands you "we have years of data, is it an asset," ask yourself: what would it cost me to be wrong about that, in each direction? Name the hidden, expensive error, optimize against it, and the rest of the answer follows.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Concealment flip: after a detailed note became evidence in a lawsuit, advisors quietly started writing vaguer, more defensive notes firmwide, with nobody deciding this on purpose.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Everett Sowande, who leads data and compliance at Averlynn Wealth Partners and audited the archive before any model touched it.
3 · THE HABIT
What did advisors stop doing after year six, because it started feeling risky?
Tap to flip
ANSWER
Writing specific, detailed meeting notes. Average note length dropped from 340 words to 61 after a detailed note became evidence in a dispute.
4 · THE FLIP, IN THIS STORY
What's the two setting switch here?
Tap to flip
ANSWER
Writing a specific, honest note versus writing a vague, defensive one. No middle setting once specificity started to feel like legal exposure.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Removing the required "recommendation outcome" dropdown in year three, to cut a click from the meeting-close screen.
6 · THE NUMBER
Fill in the blank: by year ten, only ___ % of Averlynn's notes had a usable structured outcome.
Tap to flip
ANSWER
14 percent, well below a reasonable 40 percent bar for calling the archive ready to use.
7 · THE REPLAY
Same ten years, the dropdown kept mandatory and specificity protected instead of penalized. What changes?
Tap to flip
ANSWER
A sample audit finds close to 80% of records usable instead of 14%, and the copilot project starts from a real asset instead of a pile of good intentions.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Copperlatch Veterinary's ten years of treatment charts. The flip is delegation: senior vets reclaimed detailed charting once tech-written charts stopped recording the reasoning behind a treatment choice.

Check yourself Score: 0 / 0

Multiple choice
1. According to this answer, what should you check FIRST before treating an old document archive as an asset?
  • A. How many total records exist.
  • B. Whether you actually have the right to use the archive this way.
  • C. How much it would cost to store it in the cloud.
  • D. Whether competitors have a similar archive.
Show hint
Look at the priority list's first bullet.
Show answer
B. A consent or recordkeeping violation is the one mistake here you cannot undo once a model has trained on the data.
True or false
2. True or false: this answer concludes that Averlynn's archive should never be used for any AI project.
  • True
  • False
Show hint
Look at "kill criteria" in the PICK recap.
Show answer
False. The position can flip if a compliance clearance and a real structure audit both check out. It's a testable claim, not a permanent no.
Fill in the blank
3. Fill in the blank: average note length dropped from 340 words to ___ words after the year-six lawsuit.
Show hint
Look at "prove it with the compressed failure" in the walkthrough.
Show answer
61 words. Notes went from real clinical detail to boilerplate almost overnight once specificity started to feel risky.
Short answer, where it wouldn't matter
4. Name a type of record on the same CRM where this answer says you would NOT need to force detailed structure, and say why.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: Informal internal scheduling notes, like "moved to Thursday, client running late." Nobody was ever going to train a client-facing model on those, so structuring them adds cost with no real benefit.
Short answer, apply it yourself
5. Think of an old archive your team or workplace has (emails, tickets, reports). What's one honest question you'd need answered before calling it an asset?
Show hint
Think about whether the records were written to be truthful, or written to be defensible.
Show answer
Model answer: Old customer support tickets: were agents writing honest root-cause notes, or closing tickets fast with a generic reason code just to hit a resolution-time target?
Short answer, work the number
6. If the "ready to use" bar is 40% structured and the archive dropped below that around year four, roughly how many years did the archive sit below that bar before anyone audited it in year ten?
Show hint
Look at the line chart marking the 40% threshold against the year-ten audit.
Show answer
Model answer: About six years, from roughly year four to year ten, all spent assuming the archive was fine because nobody had checked.
Before you close the answer
Why this works
Tests whether you'll take "we have years of data" at face value, or commit to a real position on what actually makes an archive usable, and name what would change your mind.
Follow-up traps
"Isn't this just being overly cautious about data you already own?" Response: no, the position isn't "never use it," it's "prove it first," which is a cheap sample audit, not a multi-year delay.

"Couldn't you just filter out the vague, post-lawsuit notes and use the rest?" Response: yes, and that's closer to the right move, treating the archive as a mixed asset with a usable slice, rather than either a blanket yes or a blanket no.
If pressed
The sample audit Everett ran used a stratified pull, twenty notes per year rather than two hundred random ones, specifically so no single busy year could hide a bad one, which is what surfaced the year-six inflection point so cleanly.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more