InterviewAdvancedResponsible AI & Advanced Practice / Compliance and legal partnership / #20

Tell me how you would prepare a product for a compliance regime that does not exist yet.

BOUND the product is Haven Companion, an AI app people talk to for emotional support and everyday check-ins

Okay, so here's the honest version of this one. Haven Companion is a chat app, you talk to it in the evening about your day, and it talks back like it actually cares. Baran Aydin is the PM on it. A colleague from trust and safety asked him a plain question in a hallway one day, and he didn't have a real answer either.

The direct answer
Here's what I'd actually do: I wouldn't try to guess the exact rule. I'd build the thing that's good product design anyway, a record of every message that's timestamped, versioned, correctable, and reviewable by a person, because that's exactly what any plausible future regime is going to ask for, whatever it ends up called. You don't need to know the law to know what "we can explain what happened" requires.
Do this, in order
  1. Give every message a real, permanent, timestamped record, not just the latest state of the conversation.Why: you can't produce a history later for something you never actually kept.
  2. Let people correct one message instead of restarting the whole conversation.Why: this is the part I'd actually build first, it's good for users today and it's the exact shape a future audit trail needs.
  3. Write consent language that says clearly what happens to the data, now, before any regulator makes you.Why: vague consent is cheap to write and expensive to redo once real rules exist and you have to explain the gap.
  4. Keep the human escalation path you probably already have for risk content, and make sure it's logged the same way.Why: don't rebuild what's already working, just make sure it leaves the same kind of record everything else does.
  5. Figure out, honestly, whether a future rule might apply retroactively, and size your estimate around that question.Why: it's the one assumption that changes the cost the most, by a wide margin.

How to answer this, stage by stage

Seven stages, and honestly, this is the one where sounding like a person matters more than sounding like a slide deck.

Stage 1
Ground it in a real product, fast
Say it like this
"Let's say it's Haven Companion, an app people talk to about their day. That's the one I'll walk through."
Why this works
A question about "a future regime" needs one real product under it immediately, or it stays a philosophy discussion.
Stage 2
Say your structure out loud
Say it like this
"I'll use BOUND for this. Break down what 'ready' even means, own my assumptions, give a range, sanity-check it, and say what swings it most."
Why this works
Signals this is an estimate with real work shown, not a guess wearing a confident voice.
Stage 3
Reframe the question
Say it like this
"Look, you can't prepare for a specific law that doesn't exist yet. What you can do is build the stuff that's good regardless, a clean record, real consent, a way to fix a mistake, and hope that whatever the rule turns out to be, you're already close."
Why this works
Stops the answer from turning into a guessing game about future legislation nobody can actually predict.
Stage 4
Give the one decision
Say it like this
"I'd build a per-message record that's timestamped, versioned, and correctable, with a real human-review path. That's the single piece of infrastructure that covers the most ground no matter what the actual rule ends up saying."
Why this works
Matches the direct answer exactly. This is the sentence someone should walk away remembering.
Stage 5
Show the arithmetic
Say it like this
"Roughly: three weeks for the per-message record structure, three for the correction feature, two for consent language, one for tying the escalation path into the same log. Call it eight weeks typical, four if a lot of it overlaps with safety work we're already doing, fourteen if the rule ends up wanting something we haven't guessed at."
Why this works
Shows the actual build-up instead of a single made-up number with no visible reasoning behind it.
Stage 6
Sanity-check it
Say it like this
"Companies that waited for GDPR to actually exist before building consent and access tooling took quarters to retrofit it. Eight weeks now, before anything's even been written into law, is cheap by comparison."
Why this works
Anchors the estimate against something the listener already has a feel for.
Stage 7
Name the direction, and close
Say it like this
"The one thing that changes this the most is whether a future rule applies retroactively, to conversations we already have. If it does, the estimate roughly triples, because old sessions were never captured at this level of detail."
Why this works
Names exactly what a good estimator flags: the single assumption a follow-up question would actually challenge.

Let's learn

Okay, so here's the plain version. Haven Companion is an app you talk to, and people use it in the evening to decompress, to talk through their day, sometimes about genuinely hard stuff.

For a long while, a conversation in the app was one long, continuous thread. That felt right, since a companion is supposed to remember things, and remembering means the thread just keeps growing.

Knowledge spark: what's a "future-proof" audit trail, really? Not a guess at a specific law. It's just a record with four things: who said what, when, whether it was later corrected, and whether a person could review it if asked. Build that, and you're covered for a lot of rules you haven't even seen written yet.

The catch: if the AI said something a little off, a little tone-deaf, in the middle of a genuinely personal conversation, there was no way to fix just that one moment. Your only real option was closing the app and starting over, which meant retyping everything you'd already said that night.

Total time to reach basic future-readiness, built up from four pieces
8 weeks 0 Escalation logging: 1wk Correction feature: 3wk Record structure: 3wk Consent language: 2wk Total: 8 weeks
Notice the correction feature and the record structure are the two biggest pieces, and they're also the two that help users today, not just some future audit.

At its worst: someone has a rough night, the AI misreads them at exactly the wrong moment, and the only path forward is to erase three weeks of context and start the whole relationship with the app over from scratch.

The decision I would take back We built every session as one continuous thread with no way to correct or annotate a single past message, only a full restart. That made sense early on, when conversations were short and a misstep barely registered. It stopped making sense once conversations got long and personal enough that one bad moment could cost someone weeks of context.

What I would leave alone: the crisis-escalation path for genuinely risky content already works and is already logged reasonably well. I wouldn't rebuild that. I'd just make sure it feeds into the same record as everything else.

We didn't just lose a message. We taught people that the only way to fix a small thing was to throw away the whole conversation.

The lesson: good record-keeping and good compliance readiness turn out to be almost the same list, most of the time. You don't need the law written yet to know what "we can show our work" requires.

Now here is the same thing as a story

Short version's above, that's what you'd say in the room. Here's the actual moment that got Baran there.

Noor Rehman opens Haven Companion most nights around ten, after her shift, and talks through whatever the day handed her. She's good at this, she says exactly what she means, first try, most of the time.

For months, if the AI missed her point, she'd just correct it in the same breath. "That's not quite what I meant, I meant X." And for a long while, that worked fine.

Hand sketched flow diagram titled The old repair path. Five boxes: AI misreads her, she notices, only option new session highlighted, three weeks of context gone, starts over.
Five steps, and the third one was never really a choice. It was the only door in the room.

Then a few nights where her correction just didn't land right, the conversation had moved on too far by the time she caught it, started to change something quieter in her. She stopped bothering to correct mid-conversation. She just started closing the app and reopening a clean session whenever a response felt off, retyping her whole day from the top.

Nobody on the team saw this as a problem. On a dashboard, it looked like more sessions, which looked like more engagement.

Hand sketched comparison diagram titled Retype it all vs fix one line. Left panel, a question mark box icon labeled Old design, caption closes app retypes her whole day. Right panel, a document icon labeled New design, caption taps one message gives one correction.
One of these looks like engagement on a chart. It's actually someone giving up on being understood.

What actually surfaced it was a colleague's remark, not a data point. Someone from trust and safety asked Baran, almost in passing, what the team would even show a regulator if one ever asked for a specific user's session history. Baran realized he didn't have a clean answer, because the sessions themselves kept getting thrown away and rebuilt from scratch by the users, not the system.

Hand sketched labeled parts diagram titled What a future ready record needs. Center document icon labeled One Turn, with four callouts: timestamp version, correction history, consent flag, human reviewable.
Four things a single message needs. Haven's messages, at the time, reliably had none of them.

The fix wasn't built as a compliance project. It was built because Noor, and users like her, deserved a way to correct one line without losing everything around it. That feature happens to produce exactly the kind of structured, timestamped, correctable record any future rule is likely to ask for.

Hand sketched timeline titled The range, low to high. Three milestones: low 4 weeks overlaps existing safety work, typical 8 weeks our best guess highlighted, high 14 weeks novel disclosure rules.
Eight weeks, give or take. Cheap, next to what a real retrofit later would have cost.

With it live, the same rough night now goes differently: Noor taps the one message that missed her, types a one-line correction, and keeps going, three weeks of context intact instead of gone.

I let sessions be one continuous, uncorrectable thread because that felt like the natural shape for a companion app, remembering everything. It took a colleague's offhand question, not an incident, to see that "remembering everything" and "being able to explain any of it later" were never actually the same promise.

BOUND, said out loudNot a spreadsheet. BOUND is what keeps an estimate about the future honest instead of just confident-sounding.

B
Break it down. The pieces of "ready."
Per-message record structure, plus a correction feature, plus consent language, plus tying escalation logging into the same record.
Says what "prepared" actually means before guessing at a number for it.
O
Own numbers. Each one sourced.
Three weeks for the record structure, three for correction, two for consent, one for escalation logging.
Every number ties back to a real, buildable piece, not a vibe.
U
Use a range. Not one guess.
Eight weeks typical, four weeks best case, fourteen weeks worst case if the rule wants something genuinely novel.
A single number here would claim a confidence nobody actually has about future law.
N
Nail the sanity check.
Companies that waited for GDPR to exist before retrofitting consent tooling took quarters, not weeks.
Compares the estimate to something the listener already has a real feel for.
D
Direction. What swings it most.
Whether a future rule applies retroactively to conversations that already happened, which could roughly triple the real cost.
The hardest step, and the one that turns a rough guess into a real, defensible plan.
Hand sketched icon list titled The assumptions behind the estimate. Four items: a document icon labeled retention length assume multi year unknown regime, a gauge icon labeled human reviewable transcript on request, a person icon labeled escalation path for risk content mostly built, a scale icon labeled consent language clear and specific.
Four assumptions, stated out loud, so anyone can push back on the one they disagree with.
Which assumption swings the estimate most, if it's wrong
0 16 weeks Retroactive requirement +16wk Novel disclosure format +5wk Consent-language scope +2wk Escalation-path change +1wk
Retroactive requirement isn't just the biggest bar here, it's more than double every other assumption combined.

The recap, one line per letter: break it down is the four real pieces of readiness, own numbers is each one's sourced estimate, use a range is four to fourteen weeks around an eight-week typical case, nail the sanity check is the GDPR-retrofit comparison, and direction is retroactive scope as the single biggest swing factor.

And if you want to be sure it really works, try it somewhere elseSame five letters, a dev-tools security scanner instead of a companion app. This time the swing factor isn't retroactive scope at all, it's who's actually liable.

Fenwick Security runs an AI tool that scans customer codebases for vulnerabilities before a release ships. Elina Saarinen manages that scanning feature.

Mapped onto BOUND: break it down is readiness equals a versioned record of every scan result, plus a clear statement of what the tool did and didn't check, plus a documented human-review path for anything the tool flags as high severity. Own numbers is two weeks for scan-result versioning, one week for a plain-language "here's what we checked" disclosure, three weeks for the human-review workflow. Use a range is four weeks typical, two weeks best case, ten weeks worst case if a future software-liability rule requires proving exactly which known vulnerabilities the tool was capable of catching at scan time. Nail the sanity check is comparing it to how long financial firms took to build model-risk documentation after banking regulators started asking for it, usually months, not weeks. Direction here isn't about old data at all, it's about liability: whether a future rule holds the tool vendor responsible for what it missed, or only the customer who shipped anyway, and that single question changes what "ready" even means far more than any technical build item does.

Hand sketched icon list reused to represent the assumptions behind a security scanner's own future readiness estimate.
A different product, a different swing factor, and the same habit of naming the assumption before betting on it.

Swap the trigger and it still runs.
Speed: an interviewer caps you at a minute. Say "build the record-keeping that's good product design anyway, and name the one assumption that could triple the cost," and stop there.
Cost: leadership says there's no budget for anything not legally required yet. Build just the correction feature first, since it pays for itself in user trust today, regardless of what any future rule ends up saying.
The model gets better, for real: if Haven's conversational quality improves a lot, that's still no reason to skip this. A better model that still can't show what it said, when, and whether it was corrected has the same readiness gap as a worse one.

Where people run it wrong.
They try to guess the actual text of a law that doesn't exist yet, and build for the wrong specifics.
They treat "we don't know the rule" as a reason to build nothing, instead of building the parts that are good regardless.
They estimate a single number with no range, which claims a confidence nobody preparing for an unknown regime actually has.

How to use it live. When this question comes up, don't try to predict the law. Ask what "we can explain what happened here" would require, and build toward that instead. It covers you no matter what the rule ends up being called.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Input flip: Noor stops correcting the AI naturally, mid-conversation, and starts restarting whole sessions instead, changing how she feeds the system rather than checking it more.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Baran Aydin, the PM on Haven Companion, and Noor Rehman, a nightly user who's good at saying exactly what she means.
3 · THE HABIT
What did Noor stop doing because correcting mid-session stopped feeling worth it?
Tap to flip
ANSWER
She stopped correcting the AI in the same breath and started closing the app and reopening a fresh session instead, whenever a response felt off.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
Correcting one line in place versus retyping the whole conversation from scratch. There was no in-between once the old repair path stopped feeling worth using.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Building every session as one continuous thread with no way to correct a single past message, only a full restart, which made sense while conversations were short.
6 · THE NUMBER
Fill in the blank: the typical estimate to reach basic future-readiness was about ___ weeks.
Tap to flip
ANSWER
Eight weeks. The range ran from four weeks best case to fourteen weeks worst case.
7 · THE REPLAY
Same rough-night moment, redesigned app. What changes?
Tap to flip
ANSWER
Noor taps the one message that missed her, gives a one-line correction, and keeps going, three weeks of context intact instead of gone.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the swing factor there?
Tap to flip
ANSWER
Fenwick Security's code scanner. There, the swing factor isn't retroactive scope, it's whether a future rule holds the vendor liable for what the tool missed.

Check yourself Score: 0 / 0

Multiple choice
1. What's the actual strategy this answer proposes for preparing for a regime that doesn't exist yet?
  • A. Wait until the specific law is passed, then build exactly what it requires.
  • B. Build the record-keeping and correction infrastructure that's good product design anyway, since it covers most plausible future rules.
  • C. Hire a lawyer to predict the exact text of the future law.
  • D. Do nothing until a competitor gets fined first.
Show hint
Look at the direct answer and stage 3's reframe.
Show answer
B. You can't predict a specific unwritten law, but a clean, correctable, human-reviewable record covers a very wide range of plausible future rules.
True or false
2. True or false: the correction feature in this story was originally built as a compliance project.
  • True
  • False
Show hint
Look at the paragraph right before "The fix wasn't built as a compliance project."
Show answer
False. It was built to give users a way to correct one line without losing everything, and it happened to double as future-readiness infrastructure.
Fill in the blank
3. Fill in the blank: the single assumption that swings this estimate the most is whether a future rule applies ___ to conversations that already happened.
Show hint
Look at the D step, direction, and the horizontal bar chart of assumptions.
Show answer
Retroactively. A retroactive requirement roughly triples the estimate, since old sessions were never captured at this level of detail.
Short answer, where it wouldn't matter
4. Name a part of Haven Companion's system that this answer says doesn't need to be rebuilt.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The crisis-escalation path for risky content already works reasonably well. It just needs to feed into the same record as everything else.
Short answer, apply it yourself
5. Pick a product you use that handles something sensitive about you. If a new rule showed up tomorrow asking "explain exactly what happened in this interaction," could it?
Show hint
Think about an app that stores health data, financial data, or personal messages.
Show answer
Model answer: Most people realize the honest answer is "probably not cleanly," which is exactly the gap this answer is built to close ahead of time.
Short answer, the number question
6. If the escalation-logging piece turned out to take five weeks instead of one, would the typical eight-week estimate still hold? Why or why not?
Show hint
Look at the stacked bar chart's build-up and which piece is smallest.
Show answer
Model answer: No, the total would rise to about twelve weeks, since the build-up is additive. But it still wouldn't change which single assumption, retroactive scope, swings the estimate most.
Before you close the answer
Why this works
Tests whether you can turn a vague, future-facing question into a real, broken-down estimate instead of either a guess or a shrug, and whether you know that good compliance prep is mostly just good product design done early.
Follow-up traps
"Isn't this just guessing what the law will say anyway?" Response: no, it's building the parts, a clean record, real consent, a correction path, that hold up under almost any plausible version of that law, not a bet on its exact wording.

"What if the eventual rule wants something totally different from what you built?" Response: then the fourteen-week worst case covers exactly that, and the gap is still smaller than starting from nothing once the rule actually lands.
If pressed
Haven's real correction feature stores both the original message and the correction as separate, linked records rather than overwriting anything, so a later reviewer can see exactly what was said, what was flagged, and what changed, in that order.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more