Artifact critiqueIntermediateResponsible AI & Advanced Practice / Compliance and legal partnership / #16
Critique a product decision that treats compliance as a launch-day task.
GUARD the artifact is Thistle Home's Launch Readiness Checklist, used before every AI feature ships
Here is what happens the first time a checklist item turns out to be decoration. Thistle Home matches customers with contractors, plumbers, electricians, HVAC technicians, using a model that also checks whether a contractor's license covers the job. Renata Alcaraz is the product manager who pulled the team's launch checklist apart after a competitor marketplace got fined over a bad match.
The artifact: Launch Readiness Checklist v3, item 9 of 9
9. Compliance review ✓ (completed day of launch). Owner: none listed. Criteria: none listed.
The direct answer
This checklist item is a rubber stamp, not a review. A real compliance check needs a named owner, written pass or fail criteria, and enough lead time before code freeze to actually block the launch if something's wrong. Put it there, before freeze, not on ship day when nobody left in the room can say no.
Do this, in order
Move compliance review to before code freeze, not the day of ship.Why: a review that happens after the code is already locked in has no real power to change anything.
Name an owner and write pass or fail criteria for the review itself.Why: an unowned checkbox with no criteria isn't a review, it's a formality nobody can be held to.
Give the reviewer actual power to block the freeze, not just flag a concern.Why: a concern raised with no power to stop anything just becomes a note nobody acts on.
Track how often concerns get raised late, and make it safe to raise them.Why: if raising a concern always means being blamed for the delay, people quietly stop raising them.
Leave the other eight checklist items exactly as they are.Why: the problem isn't the whole checklist, it's the one item that was never really a check.
How to answer this, stage by stage
Six stages. A critique question wants you to name the exact gap, not just say "this is bad."
Stage 1
Name what's actually in front of you
Say it like this
"This is a launch checklist where compliance review is the last of nine items, checked off the day of launch, with no owner and no written criteria next to it."
Why this works
Pins the critique to the exact document in front of you, not a general complaint about "not caring about compliance."
Stage 2
Say your structure out loud
Say it like this
"I'll use GUARD. Name who's exposed, where the harm lands unevenly, who can actually stop it, what I'd change, and how I'd catch this pattern again."
Why this works
Signals a structured critique instead of a gut reaction.
Stage 3
Name both people
Say it like this
"There's the VP who can order a rollback after something's already gone live, and there's the contractor who gets matched to a job needing a license they don't currently hold, with no way to flag that before it happens."
Why this works
GUARD's core move: the operator with power and the subject with none, on the same page.
Stage 4
Give the one decision
Say it like this
"Move the review before code freeze, name an owner, write real criteria, and give it the power to actually block the launch."
Why this works
Matches the direct answer exactly, so it's clear what you'd change, not just what's wrong.
Stage 5
Prove it with what actually happened
Say it like this
"A contractor gets matched to an electrical job. Their license had lapsed in that state two months earlier. The customer finds out, not the company, and files a complaint. Nobody inside Thistle caught it, because nobody was actually looking before ship."
Why this works
A compressed, real scenario proving the checklist item's failure under pressure.
Stage 6
Close on one line
Say it like this
"A checklist item with no owner and no criteria isn't a safeguard, it's a place to write the word 'done' and move on."
Why this works
Restates the critique in one breath, the sentence an interviewer would want to write down.
Let's learn
Say a home-services marketplace builds a model that matches customers with contractors, and also checks whether the contractor's license actually covers the job being booked.
For a long stretch, the team's nine-item launch checklist looked fine on paper. Every AI feature shipped with all nine boxes checked, including "Compliance review." Nobody in leadership had reason to look closer, because the box was always checked.
Knowledge spark: what makes a review real, versus decorative?
A real review has three things: someone specific who owns it, a written standard for what pass or fail actually means, and enough power to stop the thing it's reviewing. Miss any one of the three and a "review" is just a word on a checklist.
What actually happened, quietly, over about a dozen launches: engineers who used to flag last-minute regulatory concerns in the team channel learned that doing so only got them blamed for holding up a date already promised to the board. So they stopped raising them early, and concerns started surfacing only after a customer complained.
Percent of AI-feature launches with compliance sign-off completed before code freeze
For three straight quarters, "reviewed" almost never meant "reviewed before it could still be stopped."
At its worst: a contractor gets matched to an electrical job through a state where their license had quietly lapsed two months earlier. Nobody inside the company catches it. The customer does, files a complaint, and Thistle finds out the same way everyone else eventually will, from the outside.
The decision I would take back
We defaulted "compliance review" to the last item on the launch checklist, checked the same day as ship, because that was fine when the AI features shipping were low-risk and cosmetic. It stopped being fine once the model started matching people to jobs in licensed, regulated trades, where a bad match is a real safety and liability problem, not a UI nitpick.
What I would leave alone: the other eight items on the checklist, design review, load testing, rollback plans, are genuinely doing their job. The fix here is narrow, one item, not a rebuild of the whole process.
The checklist didn't fail because compliance was skipped. It failed because "reviewed" and "rubber-stamped" had quietly become the same word.
The lesson: a checklist item with no owner and no criteria isn't a safeguard sitting dormant until you need it. It's already failing, silently, every single time it gets checked.
Now here is the same thing as a story
The short version above is what you'd say defending this critique out loud. Read this one for how Marisol actually found the gap.
Renata Alcaraz had run product at Thistle Home for three years by the time a competitor marketplace, in a different city, got fined after matching a customer with an unlicensed electrician through its own AI system. It hadn't happened to Thistle. It happened to a company just like it, and that was enough to make leadership ask Marisol to check their own process.
Five steps, and the third one, where a real gate should have sat, was never actually built.
She pulled the last dozen launch checklists and found the same pattern on every single one: item 9, compliance review, checked, dated the same day as ship, no name next to it, no written criteria anywhere in the ticket.
What she actually found, across 12 launches
Owner listed: 1 of 12
Written pass/fail criteria: 0 of 12
Review completed before code freeze: 2 of 12
She talked to the engineers who'd flagged concerns in the past. More than one said some version of the same thing: raising something late got you blamed for the delay, so eventually you just stopped raising it early and let it surface however it surfaced.
One of these two people has a lever. It just happens to be the one furthest from where the actual harm lands.
She sorted the people exposed by the gap onto one page: who has the power to delay a launch, and who actually carries the risk if the match is wrong.
The people with the least power to slow anything down were exactly the ones carrying the most risk.
She rebuilt the one broken item, not the whole checklist, around what an actual review requires.
Four parts. The old checklist item had reliably had exactly none of them.
Same feature, same team. The only thing that moved was when someone with real power actually looked.
With the fix in place, a feature can't enter the freeze branch without a signed, criteria-based review attached to its ticket, six days before ship on average now instead of zero.
Four numbers, and the first two are the pair that would have caught the old pattern within a single quarter.
The old checklist asked whether nine boxes were checked. The new one asks whether the ninth box ever had the power to say no.
We put compliance last on the list because, for a long while, the features shipping genuinely were low-risk, and last felt harmless. It took a competitor's fine, not our own incident, to see that the item had quietly become a rubber stamp long before the risk caught up to it.
GUARD, applied to a checklist instead of a modelNot a policy lecture. GUARD is what forces you to name who's actually exposed by a process gap, not just that the gap exists.
G
Groups. Who's exposed.
The contractor matched to a job their license doesn't currently cover, and the customer who books them, both exposed by a gap discovered only after launch.
Names the people the checklist item was supposed to protect, not just "compliance risk" in the abstract.
U
Unequal. Where the harm lands.
Solo contractors and customers booking licensed-trade jobs carry the most exposure and have the least power to delay a launch date set months earlier.
The scramble after a bad match never lands evenly.
A
Ability to contest. The hardest step.
Once live, only a VP can order a rollback. The contractor and customer have no way to flag the gap before the match happens at all.
Asks who never gets a lever, and which past decision took it away.
R
Reduce. The concrete change.
Move the review before code freeze, name an owner, write real pass or fail criteria, and give it the power to actually block the launch.
A process decision, not a memo about "taking compliance more seriously."
D
Detect. How you'd know.
Percent of launches with pre-freeze sign-off, and days between sign-off and freeze, tracked every sprint.
Catches the pattern reopening before the next competitor's fine forces the question.
Days between compliance sign-off and code freeze, by quarter
Below the freeze line, a review has already lost its power to change anything. The whole fix is getting back above it.
The recap, one line per letter: groups is the contractor and the customer exposed by the gap, unequal is that solo contractors and licensed-trade customers carry the risk with the least power to delay anything, ability to contest is the missing pre-launch lever, reduce is a named, criteria-based, freeze-blocking review, and detect is tracking sign-off timing every sprint, not just after an incident.
And if you want to be sure it really works, try it somewhere elseSame five letters, a code-review SaaS instead of a home-services marketplace. This time the gap sits in a different kind of checklist entirely.
Driftwood CI runs an AI reviewer that flags security and licensing issues in customers' pull requests before merge. Jaan Kask manages that reviewer feature.
Mapped onto GUARD: groups is the customer whose code ships with a flagged issue nobody actually looked at, and the Driftwood engineer who owns the reviewer's rule set. Unequal is that customers on the free tier, who can't pay for a human escalation path, absorb the risk of an unreviewed flag more than paying customers with a dedicated account manager. Ability to contest is that free-tier customers have no way to ask a person to double-check a flag the AI reviewer marked "low confidence" and silently downgraded. Reduce is requiring a human glance at any flag below a stated confidence floor before it's downgraded, regardless of tier. Detect is tracking what percentage of low-confidence flags get a human look before being auto-dismissed, watched weekly instead of discovered in a support escalation.
A different product, the same four missing parts, this time behind a pull request instead of a contractor match.
Swap the trigger and it still runs.
Speed: an interviewer caps you at a minute. Say "the checklist item has no owner, no criteria, and happens after it could still stop anything, fix all three," and stop.
Cost: leadership says a full pre-freeze gate is too slow for this quarter's launch cadence. Start with the one riskiest feature category, licensed-trade matches, and expand from there.
The model gets better, for real: if the matching model's overall accuracy improves, that doesn't fix a review process with no owner and no criteria. A more accurate model with an unowned rubber stamp behind it still has an unowned rubber stamp behind it.
Where people run it wrong.
They rebuild the whole checklist when only one item was ever broken.
They add more checklist items instead of giving the existing one an owner and real power.
They treat "we do a compliance review" as true because a box gets checked, without ever asking when, by whom, or against what standard.
How to use it live. When you're handed a document like this, look for exactly three things: a named owner, written criteria, and the power to actually block something. Say out loud which of the three is missing before you say anything else.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits critiquing a compliance-as-launch-day-task decision?
Tap to flip
ANSWER
GUARD: groups, unequal, ability to contest, reduce, detect. Ability to contest is the hardest step, naming who has no lever before the harm happens.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Renata Alcaraz, product manager at Thistle Home, who audited the launch checklist after a competitor was fined for a bad contractor match.
3 · THE HABIT
What did engineers stop doing because raising concerns didn't work out well?
Tap to flip
ANSWER
They stopped flagging last-minute regulatory concerns in the open team channel, since doing so mostly got them blamed for the delay.
4 · THE FLIP
What's the two-setting switch in this story?
Tap to flip
ANSWER
Raising a concern openly, early, versus staying quiet until a customer surfaces it after launch. Once raising concerns got punished once too often, people stopped raising them at all.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Defaulting compliance review to the last checklist item, checked the day of ship, a default that was fine when features shipping were low-risk and cosmetic.
6 · THE NUMBER
Fill in the blank: only ___ of the last 12 launches had compliance review completed before code freeze.
Tap to flip
ANSWER
2 of 12. After the fix, that number became 100 percent, with an average of six days of lead time before freeze.
7 · THE REPLAY
Same licensed-trade match, redesigned checklist. What changes?
Tap to flip
ANSWER
A named reviewer checks license status against written criteria six days before freeze, with real power to block the launch, instead of a same-day checkbox nobody owns.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which product, and what's the gap there?
Tap to flip
ANSWER
Driftwood CI's AI code reviewer. There, free-tier customers have no way to get a human look at a low-confidence flag before it's silently downgraded.
Check yourself Score: 0 / 0
Multiple choice
1. What's the single biggest problem with the "Compliance review" checklist item as it was written?
A. It's the ninth item instead of the first.
B. It has no named owner, no written criteria, and happens after it could still block anything.
C. It should be removed from the checklist entirely.
D. It needs to be reviewed by a lawyer instead of an engineer.
Show hint
Look at the direct answer and the knowledge spark on what makes a review real.
Show answer
B. All three things, owner, criteria, and timing before it can still matter, are missing at once. Any one of them missing would already be a problem.
True or false
2. True or false: this critique recommends rebuilding all nine items on the launch checklist.
True
False
Show hint
Look at "what I would leave alone" and priority-list bullet 5.
Show answer
False. The other eight items are doing their job. The fix is narrow: one item, rebuilt around ownership, criteria, and real timing.
Fill in the blank
3. Fill in the blank: after the fix, compliance sign-off happens about ___ days before code freeze on average, instead of the same day as ship.
Show hint
Look at the line chart of days between sign-off and code freeze.
Show answer
Six days. Before the fix, sign-off often happened at or even after freeze, well past the point it could still change anything.
Short answer, where it wouldn't matter
4. Name a part of Thistle's launch checklist that this critique says was working fine all along.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The other eight checklist items, design review, load testing, rollback plans, and similar, were genuinely doing their job and didn't need to change.
Short answer, apply it yourself
5. Think of a checklist or sign-off step you've seen at work or school. Did it actually have an owner, written criteria, and the power to stop something, or was it closer to a rubber stamp?
Show hint
Think of an approval step that always seemed to get checked no matter what.
Show answer
Model answer: Most people can recall an approval step that always passed, which is usually a sign one of the three things, owner, criteria, or real power, was missing.
Short answer, name the reversal
6. What old decision does this critique take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Defaulting compliance review to the last checklist item, checked same-day. It made sense while the features shipping were genuinely low-risk and cosmetic.
Before you close the answer
Why this works
Tests whether you can name the exact structural gap in a document, owner, criteria, timing, rather than reacting with a general sense that "compliance should matter more," and whether you can tell a narrow fix from a rebuild nobody asked for.
Follow-up traps
"Isn't moving the review earlier just going to slow every launch down?" Response: only the launches that would have failed the review anyway, which is exactly the point, a review that never blocks anything was never really a review.
"What if the reviewer and the roadmap owner disagree about whether to ship?" Response: that's exactly why the reviewer needs real, stated power to block the freeze, not just a voice in the room that a roadmap deadline can quietly outrank.
If pressed
Thistle's real fix ties the compliance gate to a specific label on the ticket, "touches licensed trade," so only features that actually need the review are slowed down, and routine features still move at the old pace.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.