CaseAdvancedAI Opportunity & Model Strategy / Data strategy as product strategy / #8

How would you build consent and licensing into a data collection strategy from day one?

SPARKthe tablet that never asked where a photo could go

The shared tablet lives in the glovebox of six regional cars at Larkspire Leasing, one per office, and for eight months it asked field agents nothing about where a photo or a lease document could go once it was uploaded. Bram Osei runs field operations for forty agents who feed that tablet, and he is the one who has to decide what should have been built into it from day one.

The direct answer
Capture a consent and licensing tag at the exact moment a photo or document is collected, not as a policy written after the fact. Give the field agent one required choice per asset: marketing use, internal comp only, or never for training, paired with a one-tap blur suggestion for anything that looks like a tenant's personal information. Do not rely on a blanket consent clause buried in a lease and a verbal reminder to "be careful."
Do this, in order
  1. Capture a required consent and licensing tag at the point of collection, not after the fact.Why: a rule agents hear once gets reinterpreted forty different ways; a tag built into the upload flow doesn't.
  2. Limit the tag to three options, each tied to a real downstream use.Why: an agent can only apply a rule they can say back in one breath, standing in a doorway with a camera.
  3. Pair the tag with a one-tap blur suggestion for anything that looks like personal information.Why: catches signature pages and account numbers without asking every agent to become a privacy reviewer.
  4. Route anything tagged "never for training" out of the model's data pool automatically at storage.Why: a promise made at collection has to be enforced downstream, or it's just a note nobody reads.
  5. Track over-redaction, not just leaks.Why: agents blacking out real comp data to stay safe is a cost too, and it hides behind "we fixed the privacy problem."
  6. Say plainly where the standard lease consent language really is enough, like a bare exterior photo with no people or paperwork in frame.Why: shows judgment instead of tagging every single asset as if it carried equal risk.

How to answer this, stage by stage

Nobody is scoring whether you know that consent matters. They're scoring whether you can name the one moment consent actually has to be captured, and defend why after the fact is already too late.

Stage 1
Ground it in one real day-one decision
Say it like this
"Let's ground this in Larkspire Leasing, and the actual day someone has to decide what the upload flow asks a field agent, before a single photo gets taken."
Why this works
Keeps a broad governance question from turning into a policy essay.
Stage 2
State your structure
Say it like this
"I'll run this as SPARK. Situation, how data gets collected today. Payoff, the habit I want it to build. Anchor, the one decision. Risk, what breaks the first time it's wrong. Keep out, what I won't build day one."
Why this works
Signals a repeatable design method, not a list of privacy best practices recited from memory.
Stage 3
Reframe the question
Say it like this
"This isn't really 'how do you get consent.' It's 'where in the actual workflow does consent and licensing get captured,' because if the answer is 'a clause in the lease,' it was never captured anywhere an agent, or a training pipeline, can actually see it."
Why this works
This is where a strong answer separates from someone who just says "get a signed consent form."
Stage 4
Give the anchor decision directly
Say it like this
"Every photo or document gets one required tag the moment it's uploaded: marketing use, internal comp only, or never for training. Anything that looks like personal information gets a one-tap blur suggestion right there, before it's stored anywhere."
Why this works
This is the direct answer, made concrete enough that an interviewer could picture the actual screen.
Stage 5
Show what breaks the first time it's wrong
Say it like this
"Without the tag, a lease with an exposed social security number reaches the training pool the same as any exterior photo. With it, that same document gets flagged and held out automatically, no agent had to catch it by eye."
Why this works
Names the specific failure the anchor exists to survive, not a vague "there could be a leak."
Stage 6
Prove it with the compressed failure
Say it like this
"An audit sampled 150 uploads and found 22 with exposed personal information. The fix that followed wasn't a system, it was a memo saying 'be more careful.' Within six weeks, 61 of 200 new uploads were redacted so hard the comps became useless, and 9 still leaked, because nobody had actually been given a rule."
Why this works
Compresses the whole failure into the one gap a real anchor, not a memo, would have closed.
Stage 7
Name the keep-out, then close
Say it like this
"I wouldn't build a fully automatic redaction engine on day one, that's a much harder computer vision problem than this needs solved first. A tag plus a blur suggestion an agent confirms with one tap gets you almost all of the protection at a fraction of the build."
Why this works
Closes with judgment about scope, and restates the anchor in one breath.

Let's learn

Hand sketched icon list titled SPARK the five letters. Five rows: Situation, who collects data today and how, person icon. Payoff, the habit you want it to build, gauge icon. Anchor, the one decision to inspect, box icon, shown in a different color. Risk, what breaks the first time it's wrong, question mark box icon. Keep out, what you won't build day one, funnel icon.
The five letters, held up as one page. Anchor is the step this question is really testing.

Larkspire Leasing's tool reads uploaded property photos and lease paperwork to auto-generate rent comps and marketing listings. For eight months, forty agents uploaded about 1,800 property packets this way, with no tag on any single one saying what it could be used for or by whom.

Hand sketched metaphor scene titled Today, without a tag. Left, a person icon labeled Field agent, caption snaps photos, uploads everything. Right, a box icon labeled Shared tablet, caption no tag captured, ever, shown in a different color.
The agent did their job well. The tablet simply never asked the one question that mattered.
Where new uploads went wrong, a memo versus a real tag
35% 17.5% 0 30.5% 4.5% Directive only 3.5% 1% Tag plus blur
Clay bars are over-redaction, amber and green are leaks. A memo moved nothing. A tag captured at upload fixed both problems at once.

Here's the turn: the extra mistakes after the audit were never about agents caring less. They were told to "be more careful" with no system telling them what careful actually meant, so each of the forty invented their own rule, some blacking out real comp data, some barely changing anything at all.

The audit didn't create a privacy problem. It just proved one had been sitting there, uncaptured, since the day the tool launched.
The choice I would take back At onboarding, agents were told uploads were "covered under the standard lease consent language," so nobody built a per-document tag for what could be reused, or for what purpose. That made sense when a handful of agents used the tool and a manager reviewed every upload by hand. It stopped making sense once forty agents across six offices were uploading directly, with nobody reviewing in real time.

What I would leave alone: a bare exterior photo with no people, mail, or paperwork visible never needed a tag beyond what the standard lease consent already covers, and treating it like a risk just trains agents to ignore the tag everywhere else.

The lesson: consent and licensing aren't a document you write once. They're a question the collection tool has to ask, out loud, every single time something gets uploaded, or the answer defaults to "assume it's fine," which is exactly how this kind of gap opens.

Now here is the same thing as a story

The short version above is what you'd say scoping the fix in a design review. Read this one for how the gap actually widened, agent by agent, once "be more careful" was the only instruction anyone gave.

Bram Osei has run field operations at Larkspire Leasing for six years, and can usually tell from a listing's first three photos which agent shot it.

Hand sketched labeled parts diagram titled The anchor close up. A document icon at the center labeled Upload Card, with four labeled callouts around it: Consent toggle, Auto blur suggestion, Reuse purpose tag, Source agent ID.
Four things the anchor actually needs, made inspectable instead of just described in a policy.

For the tool's first eight months, agents uploaded everything they photographed on a visit, floor plans, unit conditions, sometimes a lease page left on the counter, since more documentation meant better comps and nobody had ever said otherwise. Then a regional compliance audit sampled 150 of those uploads and found 22 with a tenant's exposed personal information, a signature page here, a social security number there, sitting in the same pool the AI trained on.

Knowledge spark: why would a lease document need a different rule than a photo? A property photo is usually about the unit, not a person. A lease document is a legal record built around a specific tenant's name, income, and sometimes their government ID. The same upload button treated both the same way, which is exactly the gap a per-asset tag is built to close.

Leadership's response wasn't a new tool. It was a message to all forty agents: be more careful about what gets uploaded. With no shared definition of careful, the agents split three ways within a few weeks.

Hand sketched comparison titled The day it's wrong. Left panel, a question mark box icon labeled No tag, caption an SSN reaches the training set. Right panel, a gauge icon labeled Tag plus blur, caption flagged and held out automatically, shown in a different color.
Same document, same tenant, two very different outcomes depending entirely on whether a tag existed at the moment of upload.

Some agents kept uploading exactly as before, since the memo never reached them clearly. Others began blacking out entire pages by hand before uploading anything, including the rent concession terms and unit condition notes that made a lease useful as a real comp in the first place. A third group asked Bram directly what the rule actually was, and he didn't have an answer to give them, because there wasn't one yet.

Hand sketched flow diagram titled The upload pipeline, anchor built in, third step emphasized. Five steps left to right: Visit unit. Snap photos or docs. Tag consent and purpose. Auto blur check. Store by license tier.
The third step is the one that never existed for the first eight months, and the one every other fix depends on.

The real question was never whether Larkspire's agents cared about tenant privacy. It was whether the upload flow itself ever asked them a question they could actually answer consistently.

Weekly over-redaction rate, before and after the anchor shipped
35% 17.5% 0 wk5, peaks at 30% anchor ships Wk 1 Wk 5 Wk 10
Redaction guesswork climbed for five straight weeks under a memo. The anchor, shipped at week six, is what finally turned it around.

When the tool first launched, someone said, "the lease already covers this, let's not slow agents down with extra steps," and it sounded reasonable, since back then a manager reviewed nearly every upload by hand anyway.

Rerun the same eight months with the anchor built in from day one: every upload gets one required tag, and anything that looks like personal information gets flagged for a one-tap blur before it's ever stored. The audit never finds 22 exposed documents, because they were never routed into the shared pool. No agent ever has to guess what "be more careful" means, because the tablet already told them, every single time, before they hit upload.

What I'd tell myself, hearing that agents were blacking out real comp data by hand to stay safe: the privacy problem was never the hard part. The hard part was that nobody had ever told the tool, at the one moment it could have mattered, what any of this was actually for.

SPARK, the anchor that had to survive its own auditNot a script for slowing collection down with paperwork. SPARK is what tells you exactly which one decision a day-one consent design actually needs.

S
Situation. How data gets collected today.
Forty field agents uploading photos and lease paperwork through a shared tablet, no tag on any single asset.
One real workflow, not a general policy question, keeps the answer from staying abstract.
P
Payoff. The habit you want it to build.
Every agent making the same three-way call, consistently, in the seconds after they snap a photo, instead of forty private guesses.
The privacy protection is downstream of that habit, not a separate goal on its own.
A
Anchor. The one decision everything hangs on.
A required consent and licensing tag at upload, marketing use, internal comp only, or never for training, paired with a one-tap blur suggestion.
This is the hardest step, and the one a live answer has to name within the first minute.
R
Risk. What breaks the first time it's wrong.
An exposed social security number reaching the training pool, or the opposite failure, agents over-redacting real comp data out of fear.
Naming both failure directions, not just the leak, is what makes the anchor testable.
K
Keep out. What you won't build day one.
A fully automatic redaction engine, a much harder computer vision problem the core anchor doesn't need solved first.
Naming a deliberate boundary shows judgment, not a shortcut taken out of laziness.

The recap, one line per letter: situation is forty agents uploading through one shared, untagged tablet, payoff is one consistent habit instead of forty private guesses, anchor is the required tag plus blur suggestion at upload, risk is a leak in one direction and useless over-redaction in the other, and keep out is skipping full automatic redaction on day one.

And if you want to be sure it really works, try it somewhere elseSame five letters, a pharmacy chain instead of a leasing company. Different flip family entirely, the same day-one assumption nobody revisited.

Teodor Basu manages operations at Wickfield Pharmacy Group, where a refill-reminder tool learns from pharmacists' corrections to get better at predicting who actually needs a nudge. Mapped onto SPARK: situation is pharmacists quietly logging every override in a shared note so the team could see where the model got it wrong. Payoff is a steady stream of honest correction data feeding the next version. Anchor is exactly where Larkspire's is not: it's about visibility, not a tag, keeping override logs private and routine rather than public, since the moment leadership started publishing individual override rates on a shared leaderboard, pharmacists stopped logging honest corrections at all, worried it made them look like they didn't trust the system. Risk is the same correction data that consent depended on simply drying up. Keep out is not building any per-pharmacist performance ranking from that data, ever, since the day-one design has to assume the data source needs protecting from the company's own incentives, not just from outsiders. The flip here is concealment, not pre-editing: nobody redacted anything, they just stopped speaking up.

Hand sketched decision tree titled Which refill records train the reminder tool. Root, patient refill record. Four branches: pharmacist correction logged openly leads to keep, high weight, shown in a different color. Correction made but never logged leads to lost, unusable. Refill unchanged, no override, leads to keep, base weight. Patient opted out of research use leads to exclude entirely.
A different flip entirely: not a document getting over-redacted, but a whole stream of honest correction data going quiet.

Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "tag it at the point of collection, not after, three options, plus an auto-blur suggestion, that's the whole anchor," and stop.
Cost: there's no engineering time to build a custom tagging screen before the next release. Say so honestly, and ship the three-option tag as a simple required field first, the blur suggestion can follow once there's budget.
The data turns out to need less protection than assumed: if a legal review finds the standard lease language really does cover a specific use case, that's a real finding, not a shortcut, and it should narrow the tag's scope rather than widen it by default.

Where people run it wrong.
They write a consent policy and assume a document nobody reads at the moment of collection actually governs behavior.
They respond to a privacy incident with a reminder to "be careful" instead of a system that removes the guesswork.
They measure leaks but never over-redaction, so a fix that guts data quality looks like success on the one metric anyone's watching.

How to use it live. The moment an interviewer asks about consent and licensing, ask yourself where, physically, in the actual collection flow, someone would have to stop and make a choice. If you can't point to that moment, the consent was never really built in, it was assumed.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Pre-editing flip: with no shared rule, agents started sanitizing their own uploads, some so aggressively that the exact detail that made a document useful got stripped out.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Bram Osei, who runs field operations at Larkspire Leasing for forty agents uploading photos and lease documents.
3 · THE HABIT
What did agents stop doing after being told to "be more careful"?
Tap to flip
ANSWER
Some stopped uploading complete, usable lease documents at all, blacking out real comp data like rent concessions instead of just the personal information that actually needed protecting.
4 · THE FLIP, IN THIS STORY
What's the two setting switch here?
Tap to flip
ANSWER
Uploading documents unchanged, versus hand-redacting them so heavily the comp data inside became useless. No middle setting once agents had no shared rule to follow.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Telling agents at onboarding that uploads were already "covered under the standard lease consent language," so no per-document tag was ever built.
6 · THE NUMBER
Fill in the blank: a compliance audit sampled 150 uploads and found ___ containing exposed personal information.
Tap to flip
ANSWER
22 uploads, about 15 percent of the sample.
7 · THE REPLAY
Same eight months, the consent tag and blur suggestion built in from day one. What changes?
Tap to flip
ANSWER
No document with exposed personal information ever enters the shared pool, no agent has to guess what "be careful" means, and over-redaction never climbs past a few percent instead of peaking at 30 percent.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Wickfield Pharmacy Group's refill-reminder tool. The flip is concealment: pharmacists stopped logging honest corrections once individual override rates were published on a shared leaderboard.

Check yourself Score: 0 / 0

True or false
1. True or false: this answer argues that the actual fix Larkspire needed was a longer, clearer written privacy policy.
  • True
  • False
Show hint
Look at the direct answer and "the lesson."
Show answer
False. The fix is a required tag captured inside the actual collection workflow, not a document agents read once and reinterpret on their own.
Multiple choice
2. What does this answer say actually broke after the compliance audit, before the anchor was built?
  • A. Agents stopped caring about tenant privacy.
  • B. The AI model itself started making more mistakes.
  • C. With no shared rule, agents invented forty different private definitions of "careful."
  • D. Larkspire lost too many customers to keep collecting data at all.
Show hint
Look at "here's the turn" and the story section.
Show answer
C. A verbal directive with no system behind it let every agent guess differently, producing both over-redaction and continued leaks at the same time.
Fill in the blank
3. Fill in the blank: after the "be careful" memo, the weekly over-redaction rate climbed for five straight weeks, peaking at ___ percent before the anchor shipped.
Show hint
Look at the line chart, "weekly over-redaction rate, before and after the anchor shipped."
Show answer
30 percent. It fell to 3 percent by week 10, four weeks after the anchor's tag and blur suggestion shipped at week 6.
Short answer, where it wouldn't matter
4. Name a kind of upload where this consent tag genuinely adds nothing, and say why.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A bare exterior photo with no people, mail, or paperwork visible. The standard lease consent already covers that case, and tagging it like a risk just teaches agents to ignore the tag everywhere.
Short answer, apply it yourself
5. Think of a product you use that collects data from you or from people you work with. Where, exactly, would you want a consent choice captured, instead of buried in a terms-of-service page nobody reads?
Show hint
Look for the specific moment of collection, not a document written once, well before that moment.
Show answer
Model answer: A fitness app that shares workout data with a coach could ask, right after each session, whether that specific session can be shown to the coach or only kept private, instead of one blanket sign-up toggle set months earlier.
Short answer, work the number
6. If Larkspire had shipped the anchor at week 1 instead of week 6, roughly how many of the 200 new uploads in that window would you expect to have been over-redacted, based on the after-anchor rate?
Show hint
Use the "tag plus blur" rate from the grouped bar chart, 3.5 percent, applied to 200 uploads.
Show answer
Model answer: About 7 uploads, roughly a tenth of the 61 that were actually over-redacted under the memo-only approach.
Before you close the answer
Why this works
Tests whether you'll design consent as a real moment inside the collection workflow, or default to a written policy that nobody, human or pipeline, actually checks at the point it matters.
Follow-up traps
"Doesn't a required tag on every upload just slow agents down?" Response: it's one tap on a three-option field, not a form, and the cost of skipping it was a six-week stretch where a third of new uploads became unusable anyway.

"What if agents just pick the easiest option every time to move faster?" Response: that's exactly why the tag has to be tied to a real downstream consequence, like automatic exclusion from training, not a soft label nobody enforces.
If pressed
The blur suggestion isn't full redaction. It runs a lightweight detector for common PII shapes, an ID number pattern, a signature block, and asks the agent to confirm the blur in one tap rather than trusting an automatic system nobody double-checks.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more