Your company has ten years of unstructured documents. Is that an asset? Interrogate the claim.
Averlynn Wealth Partners is an independent advisory firm with about 240,000 client meeting notes going back ten years. Everett Sowande leads data and compliance there. When a proposal came in to fine-tune a "next-best-action" copilot on that archive, he ran an audit before anyone touched a model.
- Test the actual right to use the archive this way, before anything else.Why: a consent or recordkeeping violation is the one mistake here you cannot undo once a model has trained on it.
- Sample-audit the archive for honesty and specificity, not just for volume.Why: notes written to protect the writer, not to record the truth, teach a model the wrong lesson with total confidence.
- Check whether the archive is actually structured, or just years of free text.Why: raw prose without consistent fields is expensive to use no matter how large it is.
- Treat the volume, ten years, hundreds of thousands of records, as irrelevant until the first three checks pass.Why: a big number of unusable records is still unusable.
- Only then, decide whether to clean the archive or start fresh with a smaller, better-instrumented dataset.Why: the honest comparison is clean-and-small against dirty-and-large, not "we already have it, so let's use it."
- Revisit the archive's real value whenever a legal or regulatory event might have changed how honestly it was written.Why: what the notes actually contain can quietly change without anyone touching the file format.
How to answer this, stage by stage
Nobody is scoring whether you know that data can be valuable. They're scoring whether you can name the exact three checks that separate a real asset from a large pile of risk.
Let's learn
Here is what happens when nobody checks whether an archive that sounds valuable is actually honest.
Averlynn Wealth Partners keeps a note after every client meeting: what was discussed, what was recommended, how the client responded. Ten years of that adds up to about 240,000 records. In the first three years, the CRM had a required dropdown, recommendation accepted, declined, or deferred, and 89 percent of notes carried that structured tag alongside the free text.
A redesign in year three removed that dropdown to cut clicks from the meeting-close flow, leaving only the open text box. By year ten, only 14 percent of notes still carried anything resembling a structured, checkable outcome.
Here's the turn: the missing dropdown wasn't the real story. The real story is what happened to the free text itself. In year six, a detailed note became central evidence in a client dispute, quoted line by line in a deposition. After that, notes across the firm got shorter, vaguer, and far more careful, without anyone deciding this on purpose.
At its worst, a model trained on this archive learns two lessons nobody wanted to teach it: which advisors write the vaguest, most defensive notes, and how to sound reassuring without actually recommending anything specific, since that's exactly the pattern the post-lawsuit notes reward.
What I would leave alone: I wouldn't force detailed structure onto informal internal scheduling notes on the same CRM, since nobody was ever going to train a client-facing model on "moved to Thursday, client running late."
The lesson: an archive's age and size tell you nothing about whether it's honest. Only a sample audit tells you that, and the audit is worth running before the roadmap, not after the model.
Now here is the same thing as a story
The short version above is what you'd say defending a "don't fine-tune on this yet" recommendation in a steering committee. Read this one for how a required field's removal quietly changed what ten years of notes actually contain.
Everett has led data and compliance at Averlynn for six years. He wasn't the one who removed the dropdown. He was the one who had to explain, years later, why it mattered that someone had.
The proposal that landed on his desk was simple on paper: fine-tune a copilot on ten years of meeting notes so it could suggest a next best action for any client, using the firm's own accumulated judgment. The pitch deck's first slide said "240,000 records of proprietary advisor expertise." Nobody in that room had actually read a random sample yet.
Everett pulled two hundred notes at random, evenly spread across all ten years. The pattern by year was obvious within an afternoon. Years one through three read like real clinical judgment: "Client expressed concern about market volatility; recommended shifting 10% from equities to bonds; client accepted." Years seven through ten mostly read like: "Met with client. Reviewed portfolio. No changes needed at this time."
Everett found the year-six lawsuit in an old compliance memo, not in the CRM itself. The note quoted in the deposition had specified an exact percentage shift and an exact reason, and a client later argued that reason had been wrong. Nobody blamed the note-taking. Everyone quietly changed how they wrote notes afterward anyway.
When the dropdown was removed in year three, someone said, "let's simplify the meeting-close screen, advisors are already writing it in the notes anyway," and it sounded completely reasonable, since at the time, most of them genuinely were.
Rerun the same ten years with the dropdown kept mandatory, and a firm policy that a note's specificity is protected, not penalized, in any legal review: the archive still shows the same year-six dispute, but the notes around it stay just as detailed as the years before, because nobody ever learned that detail was a liability. A sample audit ten years later finds close to 80 percent of records usable instead of 14, and the copilot project starts from an actual asset instead of a pile of good intentions.
What I'd tell myself, reading two hundred notes that got steadily vaguer for no reason anyone had chosen on purpose: the archive didn't fail because advisors got worse at their jobs. It failed because nobody protected the one thing that made it valuable, which was advisors feeling safe enough to write down exactly what they actually did.
PICK, the call that would have caught this before the pitch deckNot a definition of what makes data valuable. PICK is what forces you to commit to a position and name what would change it.
The recap, one line per letter: position is treating the archive as raw material until proven otherwise, impact is naming both the client's hidden exposure and the team's visible delay, cost asymmetry is the roughly $2.4M hidden risk against a $180K known cost, and kill criteria is a compliance clearance plus a real structure audit, either of which would flip the call.
And if you want to be sure it really works, try it somewhere elseSame four letters, a veterinary practice instead of a wealth manager. Different flip family entirely, the same interrogated claim.
Dr. Wendeline Ashcroft runs Copperlatch Veterinary, which has ten years of handwritten treatment charts, now scanned, that a vendor wants to fine-tune into a treatment-recommendation assistant. Mapped onto PICK: position is that the charts are raw material until their consistency is proven, not an asset by page count. Impact is a bad recommendation reaching an animal whose owner has no way to catch a subtle dosing error, against the slower, visible cost of building from two years of clean digital records instead. Cost asymmetry favors caution, since a wrong dosing suggestion is the hidden, expensive error. Kill criteria is a sample audit confirming the charts consistently record the actual clinical reasoning, not just a diagnosis code.
The flip here is delegation, not concealment. As Copperlatch grew, senior vets started handing chart-writing to vet techs to save time during busy shifts, then quietly took detailed charting back for themselves once they noticed tech-written charts recorded the diagnosis and treatment but never the reasoning behind choosing one drug over another, which is exactly what a future recommendation model would need to learn from.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "it's raw material until rights, honesty, and structure are checked, not an asset by default," and stop.
Cost: no budget for a full compliance review before a pilot needs to start. Say so honestly, and run the cheap sample audit first, since that alone often settles the question.
The archive turns out to be genuinely fine, for real: if the sample audit shows strong, consistent, cleared records, that's a legitimate reason to proceed with confidence, not a shortcut being taken.
Where people run it wrong.
They treat record count as a proxy for data quality, without ever sampling the actual content.
They assume old records were captured under consent that covers a brand new use, like model training, without checking.
They discover the archive's real condition only after a model has already trained on it, instead of auditing first.
How to use it live. The moment an interviewer hands you "we have years of data, is it an asset," ask yourself: what would it cost me to be wrong about that, in each direction? Name the hidden, expensive error, optimize against it, and the rest of the answer follows.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Couldn't you just filter out the vague, post-lawsuit notes and use the rest?" Response: yes, and that's closer to the right move, treating the archive as a mixed asset with a usable slice, rather than either a blanket yes or a blanket no.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Data strategy as product strategy
- #1 Explain why the data you collect today determines the products you can build in two years.
- #2 What is a data flywheel and what are its preconditions?
- #3 Describe how you would instrument a product to generate training or eval data as a byproduct.
- #5 How do you evaluate whether proprietary data is actually a moat?
- #6 What are the product implications of not owning your own data?
- #7 Describe the difference between data volume, data quality and data relevance for AI products.