CaseAdvancedResponsible AI & Advanced Practice / Compliance and legal partnership / #19

Describe how you would respond to a regulator's inquiry about your AI feature.

TRACE the product is Birchwell Invest, a robo-advisor that rebalances client portfolios automatically

Birchwell Invest's model recommends how to rebalance a client's portfolio and can execute the trade automatically. Renata Duarte is the product manager for that rebalancing model. She got the call on a Tuesday: a state regulator wanted to know why a retired client's "conservative" account held forty percent in a single stock.

The direct answer
Don't respond with the aggregate number, respond with the timeline and the segment. Find exactly when the problem started, show which specific group of accounts it actually hit hardest, name the real cause, not a guess, and hand the regulator a dated remediation plan alongside the records they asked for. An inquiry answered with "our average looks fine" fails before it starts.
Do this, in order
  1. Build the exact timeline before saying anything else.Why: a regulator's first question is always when, and guessing at that erodes trust in everything you say after.
  2. Slice the data by client segment, not just the overall average.Why: an average that looks fine can hide one segment doing badly, and that's usually exactly the segment the regulator is asking about.
  3. Rule out a reporting or dashboard error before blaming the model.Why: a tracking glitch can look identical to a real problem, and confusing the two wastes the regulator's time and yours.
  4. Name the confirmed root cause, with the one test that proved it.Why: "we're looking into it" is not an answer a regulator can act on.
  5. Hand over a dated remediation plan alongside the records requested.Why: an inquiry wants to know it won't happen again, not just what happened once.

How to answer this, stage by stage

Seven stages. A regulator's inquiry rewards exactly the opposite instinct from a normal incident review: slower, more documented, less defensive.

Stage 1
Scope it to a real inquiry
Say it like this
"I'll answer this for Birchwell's rebalancing model, after a regulator asks why one client's conservative account held too much of a single stock."
Why this works
Grounds an abstract "regulator inquiry" question in one real complaint with a real client behind it.
Stage 2
Say your structure out loud
Say it like this
"I'll use TRACE. Timeline, recut by segment, assume nothing, cause candidates, then the one evidence test that separates them."
Why this works
Signals a real investigation, not a rehearsed apology.
Stage 3
Reframe the question
Say it like this
"This isn't really about writing a good letter back. It's about whether I can actually find the true cause before I say anything to a regulator I can't later take back."
Why this works
Moves past "communications strategy" and into the actual investigative judgment being tested.
Stage 4
Give the one decision
Say it like this
"Build the timeline first, then slice by client segment before I say anything about the aggregate. The real problem is almost never visible in the average."
Why this works
Matches the direct answer exactly. An interviewer should be able to write this down as your whole approach.
Stage 5
Prove it with the actual number
Say it like this
"Here's what I'd find. The overall drift-from-policy metric barely moved, 2.1 to 2.9 percent. But conservative-tier accounts specifically went from 2.0 to 11.4 percent, because our fixed-size daily review sample stopped covering that segment as total accounts tripled."
Why this works
A concrete, checkable number, showing exactly where the average was hiding the real story.
Stage 6
Rule out the easy explanation
Say it like this
"Before I blame the model, I'd check whether this is a dashboard labeling bug instead. In this case it wasn't, the actual recommended trades themselves, checked directly against policy limits, confirmed the drift was real."
Why this works
Shows the discipline of ruling out instrumentation before behavior, TRACE's core habit.
Stage 7
Close on one line
Say it like this
"Answer a regulator with the timeline and the segment that actually broke, not the average that makes everything look fine."
Why this works
Restates the decision in one breath, closing on the actual point instead of a vague promise to cooperate.

Let's learn

Say a robo-advisor manages tens of thousands of client accounts, rebalancing each one automatically to match the client's stated risk tolerance.

Eighteen months ago, Birchwell's compliance team reviewed a full daily sample of fifty accounts, pulled at random, against every account then under management, about forty thousand of them. That sample felt thorough at the time, since fifty was a real, meaningful slice of forty thousand.

Knowledge spark: what's portfolio drift? How far a client's actual holdings have wandered from the mix they agreed to. A conservative client who agreed to a diversified, low-risk mix but ends up holding forty percent in one stock has drifted a long way from that agreement.

Account volume more than tripled over those eighteen months, but the daily review sample stayed fixed at fifty accounts, a number nobody ever revisited. As total volume grew, that same fifty represented a shrinking slice of the whole, and conservative-tier accounts, always a minority of the total, got squeezed out of the sample almost entirely.

Portfolio drift-from-policy, by segment, over eighteen months
12% 0 2.1% 2.9% Aggregate 2.0% 11.4% Conservative tier ■ 18 months ago ■ now
The aggregate barely moved. One segment, the smallest, quietly went to nearly six times its old drift level.

At its worst: a retired client with a conservative risk profile ends up holding forty percent of her account in a single technology stock, discovers it herself while checking her statement, and files a complaint that reaches a state regulator before it ever reaches Birchwell's own compliance desk.

The decision I would take back We fixed the daily review sample at fifty accounts, regardless of total volume, because fifty felt like a meaningful, defensible number when the whole book was forty thousand accounts. It stopped being meaningful once volume tripled and that same fifty represented a much smaller, less representative slice, especially for a minority segment like conservative-tier clients.

What I would leave alone: the aggressive-tier review process didn't need to change. That segment is large enough that even a shrinking fixed sample still caught real problems there. The gap was specific to a smaller segment, not universal.

The average told us everything was fine. It just never had a way to say which one account in ten thousand it was fine for.

The lesson: a review sample that felt thorough on day one can quietly stop being thorough as volume grows, with no single moment marking when it happened.

Now here is the same thing as a story

The short version above is what you'd say drafting the regulator response. Read this one for how Renata actually traced it.

Renata Duarte has managed Birchwell's rebalancing model for three years, long enough to remember when the daily fifty-account review genuinely felt like a real check on the whole system.

There was no single bad day that broke it. Account volume climbed steadily, quarter after quarter, while the fifty-account sample stayed exactly the same size, a number set once and never revisited as the business grew around it.

Hand sketched flow diagram titled The review sample, shrinking in place. Five boxes: all accounts, 50 sampled daily, same 50 tripled base highlighted, conservative tier thins out, gap goes unseen.
Five steps, and the third one is where a fixed number quietly stopped meaning what it used to.

Then the regulator's letter arrived: a formal inquiry, prompted by a retired client's complaint, asking Birchwell to explain how a conservative account ended up forty percent concentrated in a single stock, and to produce records for a sample of similar accounts.

Hand sketched timeline titled When it started vs when it was reported. Four milestones: risk data source changed month 1 quiet highlighted, conservative tier drift begins month 2, client complains month 9, regulator inquiry month 10.
The real start was eight months before anyone outside the company noticed anything was wrong.

Renata's first move wasn't to explain. It was to rebuild the timeline properly, and then slice the data by segment instead of trusting the aggregate number the team had been watching all along.

Hand sketched comparison diagram titled Full sample vs a fixed slice. Left panel, a document icon labeled 18 months ago, caption 50 of 40,000 accounts reviewed daily. Right panel, a question mark box icon labeled Today, caption same 50 out of 130,000 accounts.
The number on the review sample never changed. What it actually represented did, completely.

Before pointing at the model, she ruled out the easier explanation: maybe the dashboard itself was miscounting drift, not the recommendations. A direct check of the actual recommended trades, bypassing the dashboard entirely, confirmed the drift was real, not a reporting artifact.

Hand sketched decision tree titled The three suspects. Root: why did conservative tier drift rise. Three branches: risk sub model drifted leads to confirmed by direct check, review sample missed it leads to contributing not root cause, dashboard reporting bug leads to ruled out.
Three named suspects, and only one of them survived a direct look at the real recommendations.

She sorted every client segment by two things: how big a share of the total book it was, and how much of the daily review sample it actually got.

Hand sketched quadrant titled Where the blind spot actually sat. Axes share of total accounts and share of daily review sample. Aggressive tier and balanced tier sit toward the upper right, well reviewed. Conservative tier sits lower left, small share and rarely reviewed.
The smallest segment got the smallest share of review, exactly the opposite of where the real risk was building.

With the root cause confirmed, direct check, not a guess, Renata assembled the response the regulator actually needed.

Hand sketched labeled parts diagram titled What the regulator actually gets. Center document icon labeled The Response, with four callouts: full timeline, segment data, root cause, remediation plan.
Four parts, and a response missing any one of them reads as either incomplete or evasive.

The old review asked whether fifty accounts, sampled the same way every day, still looked fine. The new one scales the sample with total volume and stratifies it by segment, so a minority segment can't quietly fall out of view again.

I set the review sample at fifty because it was a real, defensible number the day I set it. It took a regulator's letter, not our own dashboard, to see that a fixed number stops representing the same thing once everything around it triples.

TRACE, for a regulator's letterNot a defense. TRACE is what forces you to find the real cause before you ever put anything in writing to someone who can hold you to it.

T
Timeline. When it actually started.
A risk-data source changed quietly in month one. Conservative-tier drift began the next month. The complaint and inquiry landed eight and nine months later.
The gap between when it started and when anyone outside the company noticed is the whole story.
R
Recut. Slice by segment.
Aggregate drift barely moved, 2.1 to 2.9 percent. Conservative-tier drift specifically went from 2.0 to 11.4 percent.
The average was never lying. It just wasn't built to show one segment cratering.
A
Assume nothing. Rule out instrumentation first.
Checked whether the dashboard itself was miscounting drift before concluding the model's recommendations were actually wrong.
A tracking bug can look exactly like a real problem on a screen.
C
Cause candidates. Three named, not a guess.
The risk sub-model drifting, the review sample missing the segment, or a dashboard reporting bug.
Naming candidates, instead of jumping to the first plausible story, is what makes the next step possible.
E
Evidence test. The hardest step.
Checking the actual recommended trades directly against policy limits, bypassing the dashboard, confirmed the risk sub-model as the real cause.
The one check strong enough to separate the real cause from the two plausible-sounding decoys.
Aggregate drift metric, monthly, with the real start date marked
4% 0 Month 1: real cause begins Month 10: inquiry lands
The aggregate line barely bends. The real story was never visible in this view at all, only in the segment cut.

The recap, one line per letter: timeline is the eight-month gap between the real start and the inquiry, recut is conservative-tier drift far outpacing the aggregate, assume nothing is ruling out a dashboard bug first, cause candidates is naming three real suspects, and evidence test is checking the actual trades directly to confirm the true cause.

And if you want to be sure it really works, try it somewhere elseSame five letters, a food-delivery marketplace instead of a robo-advisor. This time the regulator asking questions is a labor regulator, not a financial one.

Cartwheel Eats runs an AI system that scores delivery drivers and adjusts how often they're offered high-paying orders. Merritt Sloane manages that scoring model.

Mapped onto TRACE: timeline is when a routing-efficiency update quietly began weighting delivery speed more heavily, three months before a labor regulator's inquiry into pay disparities. Recut is finding that drivers in one specific city, with older phone hardware that reported location less precisely, scored lower on the speed metric and got fewer high-paying offers, while the citywide average barely moved. Assume nothing is checking whether the location data itself was the problem, a hardware and reporting issue, rather than assuming the scoring model was unfairly targeting anyone. Cause candidates are the routing update, a hardware-driven data quality gap, or a genuine behavioral difference in that city's drivers. Evidence test is comparing scores for drivers with newer phones in that same city against drivers with older phones, which isolated the hardware-driven data gap as the real, confirmed cause.

Hand sketched quadrant reused to represent which driver segments were under reviewed relative to their share of the platform.
A different platform, a different regulator, and the same shape of blind spot in a different corner.

Swap the trigger and it still runs.
Speed: an interviewer caps you at a minute. Say "timeline first, then segment, rule out instrumentation, then one test to confirm the real cause," and stop.
Cost: legal wants a response sent within twenty-four hours. Send an honest interim letter naming what's confirmed so far and when the full findings will follow, rather than guessing at a cause to look responsive.
The model gets better, for real: if the rebalancing model's overall accuracy improves, that doesn't mean a review process built for a much smaller book is still adequate. A better model with a stale review process can still miss the same segment.

Where people run it wrong.
They respond with the aggregate number, which is exactly the number that made the problem invisible in the first place.
They assume the model must be the cause without checking whether a dashboard or reporting bug could look identical.
They treat the regulator's specific question as the whole scope, instead of checking whether the same gap exists elsewhere too.

How to use it live. When this question comes up, say the word "segment" out loud before you say anything about the model. That's the habit that actually catches what an average hides.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Scope flip: the review shrank from covering a meaningful slice of every account to a fixed, tiny fraction as volume tripled, and the unreviewed slice is where the real problem hid.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Renata Duarte, who has managed Birchwell Invest's rebalancing model for three years.
3 · THE HABIT
What did the review team keep doing long after it stopped working?
Tap to flip
ANSWER
Reviewing the same fixed fifty accounts a day, a number set once when the whole book was forty thousand accounts and never revisited as it tripled.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
Reviewing a genuinely representative slice of the whole book versus reviewing the same fixed number that had quietly become a much smaller, less representative fraction.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Fixing the daily review sample at fifty accounts regardless of total volume, a number that felt meaningful at forty thousand accounts and stopped being meaningful at over a hundred thousand.
6 · THE NUMBER
Fill in the blank: conservative-tier drift rose from 2.0 percent to ___ percent, while the aggregate barely moved.
Tap to flip
ANSWER
11.4 percent. The aggregate only moved from 2.1 to 2.9 percent over the same period.
7 · THE REPLAY
Same regulator inquiry, redesigned review process. What changes?
Tap to flip
ANSWER
The review sample scales with total volume and is stratified by segment, so a minority segment like conservative-tier accounts can't quietly fall out of coverage again.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the confirmed cause there?
Tap to flip
ANSWER
Cartwheel Eats' driver-scoring model. There, older phone hardware in one city caused a data-quality gap, not a genuine behavioral difference among drivers.

Check yourself Score: 0 / 0

Short answer, name the reversal
1. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Fixing the daily review sample at fifty accounts regardless of total volume. It made sense while fifty was still a meaningful slice of a forty-thousand-account book.
Multiple choice
2. Why does this answer check for a dashboard reporting bug before blaming the rebalancing model?
  • A. Dashboard bugs are always the real cause in financial systems.
  • B. A tracking or reporting error can look identical to a real behavioral problem, and confusing the two wastes everyone's time.
  • C. Regulators only care about dashboards, not models.
  • D. It's required by law to check the dashboard first.
Show hint
Look at the A step, assume nothing, in the TRACE recap.
Show answer
B. TRACE's assume-nothing step exists because instrumentation problems and real behavior problems can look exactly the same on a screen.
True or false
3. True or false: the aggregate drift metric clearly showed the problem before the regulator's inquiry arrived.
  • True
  • False
Show hint
Look at the line chart of the aggregate drift metric.
Show answer
False. The aggregate barely moved, from 2.1 to 2.9 percent. Only the segment-level cut for conservative-tier accounts revealed the real severity.
Fill in the blank
4. Fill in the blank: the real underlying cause began about ___ months before the regulator's inquiry arrived.
Show hint
Look at the timeline diagram and the T step in the recap.
Show answer
Eight or nine months. The drift began in month one or two, and the inquiry landed around month ten.
Short answer, where it wouldn't matter
5. Name a client segment where this exact review-shrinking problem did not cause real harm.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The aggressive-tier segment, since it's large enough that even a shrinking fixed sample still caught real problems there.
Short answer, apply it yourself
6. Pick a product you use yourself. What's one segment of its users that might be quietly underserved by an average-looking metric?
Show hint
Think about a minority group of users: a different device, a different language, a different plan tier.
Show answer
Model answer: Most people can name a smaller user group, like older devices or a less common language, that a healthy-looking average could easily be hiding.
Before you close the answer
Why this works
Tests whether you'll investigate before you explain, and whether you know a regulator wants a timeline, a confirmed cause, and a real fix, not reassurance dressed up as an answer.
Follow-up traps
"Why not just tell the regulator the average metric looked fine?" Response: the average is exactly what let this go unnoticed for eight months, and leading with it would look evasive the moment the regulator's own segment-level complaint contradicts it.

"How do you know the review sample gap is really the root cause and not just a contributing factor?" Response: the direct check of actual recommended trades against policy limits confirmed the model itself was recommending bad allocations, independent of whether anyone reviewed them.
If pressed
Birchwell's real remediation ties the review sample size to a formula based on total account count and stratifies it proportionally by risk tier, so no single segment can fall below a minimum review share again, regardless of how fast the business grows.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more