ConceptAdvancedResponsible AI & Advanced Practice / Internal AI tooling and enablement products / #8
What governance does an internal AI tool need that a customer product does not?
GUARD the tool: Kestrel Mill Textiles, and AskHR, its internal assistant for benefits, leave, and pay questions
Interviewer's question: "What governance does an internal AI tool need that a customer product does not?" Kestrel Mill Textiles runs a fabric mill with about six hundred floor and office employees. Grace Mbeki works as an HR generalist there.
The direct answer
An internal AI tool needs a mandatory human sign-off on anything touching pay or protected leave, not an optional review step, because the subject of a wrong answer here often doesn't know there's anything to contest. A customer product's user can complain, cancel, or leave a bad review. An employee getting a confidently wrong benefits answer from a company system usually just accepts it, especially if nobody ever tells them a human hadn't actually confirmed it.
Do this, in order
Require a human sign-off on anything affecting pay or protected leave, not advisory review.Why: the employee on the receiving end usually has no way to know the answer was ever wrong.
Show a plain "confirmed by a person" or "AI-drafted, unconfirmed" flag on every determination.Why: without it, a confident guess and a verified answer look identical to the person reading it.
Build a way to detect harm without waiting for a complaint.Why: the employees least likely to question a wrong answer are also the least likely to file one.
Never let a single query collapse a complex case into fragments.Why: a case split into isolated pieces can hide an interaction between two policies that matters only when read together.
Name who has the least power to push back, and design for them first.Why: a governance plan built around the confident, well-connected employee misses exactly who needs protecting.
How to answer this, stage by stage
Nobody is grading whether you can name "bias" as a risk. They're grading whether you can name who specifically can't push back, and what design gives them a lever.
Stage 1
Scope it to one tool, one company
Say it like this
"I'll answer this for Kestrel Mill Textiles and AskHR, the internal assistant that answers employee questions about benefits, leave, and pay."
Why this works
Keeps a governance question from turning into an abstract policy discussion.
Stage 2
Say your structure out loud
Say it like this
"I'll use GUARD. Groups affected, where harm lands unevenly, who can't contest it, the design change to reduce it, and how I'd detect misuse."
Why this works
Signals a serious method before naming a single harm, so it doesn't read as improvised concern.
Stage 3
Name both groups
Say it like this
"There's Grace, who operates AskHR and can escalate, override, or ask again. And there's the employee on the receiving end of its answer, who usually can't do any of that."
Why this works
GUARD's strongest move: naming the operator and the subject, not just "users" in general.
Stage 4
Say where harm lands unevenly
Say it like this
"A probationary night-shift worker, or someone new to the country, is the least likely to question a confident wrong answer, and the least likely to know they even could."
Why this works
Answers the real follow-up: harm from a wrong answer isn't spread evenly across employees.
Stage 5
Give the design change
Say it like this
"Any determination touching pay or protected leave needs a real human sign-off, not an optional review, plus a plain flag showing whether a person actually confirmed it."
Why this works
This is the direct answer: a concrete design change, not a policy document.
Stage 6
Say how you'd detect misuse, and close
Say it like this
"Since the harmed employee usually won't file a complaint, I'd audit a random sample of determinations against policy every quarter and track the rate anything gets overturned once someone finally does ask."
Why this works
Shows a real detection plan for harm that never generates its own support ticket.
Let's learn
AskHR answers employee questions about benefits, leave eligibility, and pay, drawing on the company handbook and the union's collective agreement.
Grace uses it daily to draft answers faster than she could by cross-referencing both documents herself. On a routine question, it's usually right, and she signs off on its draft in under a minute.
AskHR determinations later overturned on appeal, by employee tenure
A low overturn rate for new employees isn't good news. It's a sign they don't know they can ask.
The turn: the low overturn rate among new employees was read, for a while, as proof AskHR was more accurate for them. It wasn't. It meant they were less likely to know an answer could be wrong at all, let alone how to challenge one.
Both people are looking at the same kind of answer. Only one of them has a lever to pull if it's wrong.
The decision that mattered
A complex leave case, involving more than one overlapping policy, had no structured way to be reviewed as a whole. Grace could ask AskHR about the whole case, or one clause at a time, but nothing showed her the interaction between clauses either way.
At its worst: a worker's protected leave gets miscalculated because two policies interacted in a way nobody caught, and the worker, unaware anything was ever uncertain, gets flagged for unauthorized absence and terminated over a mistake they never had a chance to contest.
The fourth step is empty. That's not a missing feature, it's the actual governance gap.
What I would leave alone: routine, low-stakes questions, like how many vacation days remain this quarter, don't need this level of scrutiny. A wrong answer there is annoying, not life-altering, and gets corrected the next time someone checks their balance.
The lesson: a customer who's wronged can walk away and tell someone. An employee who's wronged by a company system often has nowhere obvious to walk to, and designing as if they do is how the harm stays invisible.
Multi-policy leave cases where the real interaction was caught, before and after the structured view
The catch rate had already fallen for months before the wrongful termination. The structured view didn't just fix one case, it fixed the slope.
Now here is the same thing as a story
The short version above is what you'd say to Kestrel Mill's HR director. Read this one for how the case actually unfolded.
Grace Mbeki has worked HR at Kestrel Mill for three years. She's good at spotting when a leave request touches more than one policy, the kind of case that takes real judgment, not a form.
For most of her time using AskHR, she asked it about a whole case at once, the same way she'd think through it herself. It handled routine cases well. Then came a case with two overlapping leave provisions, medical leave and a separate accommodation policy, and AskHR's answer, drafted for the whole case at once, missed how the two interacted.
Knowledge spark: why can two correct-sounding policy answers add up to a wrong one?
Some rules only make sense combined. A leave policy might say "up to twelve weeks," and a separate accommodation policy might extend that under specific conditions. Answered alone, each rule sounds complete. Answered together, the true entitlement can be longer than either rule states by itself.
Grace caught that particular miss by luck, cross-referencing the two policies herself out of habit. Rattled, she changed how she used AskHR afterward: instead of asking about a whole complex case, she started asking about one clause at a time, wanting to double-check each piece carefully.
The second milestone felt like extra caution. It's exactly what let the third one happen.
Months later, another complex case came through, this time while Grace was managing a busy week. She asked AskHR about each clause separately, confirmed each piece looked right, and never re-assembled the full picture the way she had, by instinct, the first time. The interaction between two policies got missed again, quietly, and this time nobody caught it before the decision went out.
The worker didn't lose an argument about their leave. They never knew there was one to have. The termination notice cited unauthorized absence, and nothing in it said a policy interaction, not the worker, was the actual mistake.
Here's the decision I'd take back: never building a way to handle a complex case in structured pieces safely. The all-or-nothing choice, ask about everything at once or ask about one clause at a time, forced Grace into fragmenting her own view the moment a case got too complex to hold in one query. That made sense when most cases were simple enough not to need it. It stopped making sense the day a case wasn't.
Replayed with a structured breakdown that shows every clause and their interactions on one screen, confirmed by a person before anything is final: the same overlapping-policy case gets flagged automatically, since the design forces the interaction to be visible, not something Grace has to reconstruct from memory. The correct, longer leave period gets approved, and the termination never happens.
I let the query stay all-or-nothing because building structured, cross-referenced views felt like solving a problem nobody had reported yet. It took a wrongful termination to see that nobody reporting it wasn't the same as nobody being harmed by it.
GUARD, run against one leave caseNot a compliance checklist. GUARD is what tells you who has no lever to pull when the answer is wrong.
G
Groups. Who is affected.
Grace, who operates AskHR, and the employee whose leave or pay is decided by its answer.
Names the operator and the subject, not a vague "users" category.
U
Unequal. Where harm lands unevenly.
Probationary, newer, and less-connected employees are least likely to know a wrong answer can be challenged at all.
Backs the claim with a real number: a 2 percent overturn rate for new hires versus 18 percent for veterans.
A
Ability to contest. The strongest move.
A worker terminated over a missed policy interaction never knew there was anything to appeal in the first place.
The hardest step: naming who has no lever, not just that harm exists.
R
Reduce. The design change.
Mandatory human sign-off on anything touching pay or protected leave, with a plain confirmed-or-not flag on every determination.
A concrete product decision, not a policy memo nobody reads.
D
Detect. How you'd know without a complaint.
Random quarterly audits against policy, and tracking the overturn rate, since the harmed employee often won't be the one to report it.
A real detection plan for harm with no built-in support-ticket signal.
Two of these four callouts are missing entirely. Both of them are governance, not features.
All three exist specifically because the harmed party often never files anything themselves.
The bottom-left corner is exactly who a governance plan has to be built around, not an edge case to handle later.
The recap, one line per letter: groups is Grace and the employee, unequal is newer and less-connected employees carrying more risk, ability to contest is a worker who never knew to appeal, reduce is mandatory sign-off with a visible flag, and detect is auditing without waiting for a complaint.
And if you want to be sure it really works, try it somewhere elseSame five letters, a community bank's loan-modification desk instead of a textile mill's HR office. Nothing about the two jobs is alike.
Corrigan Trust Bank gives loan officers an internal assistant that evaluates whether a struggling borrower qualifies for a loan modification. Malik Osei is a loan officer there.
Mapped onto GUARD: groups are Malik, who operates the tool and can escalate or override it, and the borrower facing default, who receives its determination. Unequal: borrowers with less financial literacy or English fluency are least likely to know a modification denial can be appealed at all. Ability to contest is the strongest parallel: a borrower who's denied and never told a human hadn't confirmed the decision has no reason to think there's anything left to ask for. Reduce is the same design change: mandatory human sign-off on any denial, with a visible confirmed-or-not flag. Detect is structurally identical too: random audits of denials against lending policy, since a borrower who quietly accepts a wrong denial and walks away never generates a complaint about it.
A different shape of picture than Section 2 used: a branching path instead of a labeled document. The missing branch is the same kind of gap either way.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "require human sign-off on anything touching pay or leave, flag whether a person confirmed it, and audit without waiting for a complaint," and stop.
Cost: if a full structured cross-reference view is too expensive to build immediately, start by simply forcing any multi-policy case to route to a human, no exceptions, until the tooling catches up.
The model gets better, for real: even as AskHR's accuracy improves overall, the subject's inability to contest a wrong answer doesn't improve with it, so the governance requirement doesn't shrink just because errors get rarer.
Where people run it wrong.
They treat a low complaint rate as proof a tool is working well for everyone, without checking whether the quiet group even knows how to complain.
They build governance around the confident, well-connected employee who will ask questions, and miss who actually needs the protection.
They add a review step that's optional or easy to skip under time pressure, instead of one that's mandatory for anything touching pay or protected leave.
How to use it live. When someone asks what governance an internal tool needs, ask yourself first: if this tool is wrong about someone, will that person even know to say so? If the honest answer is no, that's exactly where the mandatory human check has to go.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits "what governance does an internal AI tool need that a customer product does not"?
Tap to flip
ANSWER
GUARD: groups, unequal, ability to contest, reduce, detect. Ability to contest is the strongest step here.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Grace Mbeki, an HR generalist at Kestrel Mill Textiles, good at spotting when a leave case touches more than one policy.
3 · THE HABIT
What did Grace stop doing after her first near miss with AskHR?
Tap to flip
ANSWER
She stopped asking about a whole complex case at once and started asking about one clause at a time, losing the full picture in the process.
4 · THE FLIP
What's the two-setting switch in this story?
Tap to flip
ANSWER
Reviewing a whole case together, versus reviewing it clause by clause and never reassembling the interactions between the pieces.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Never building a structured way to review a complex, multi-policy case safely, which forced an all-or-nothing choice the moment a case got too complex.
6 · THE NUMBER
Fill in the blank: new employees had an overturn rate of 2 percent, versus ___ percent for employees with five or more years of tenure.
Tap to flip
ANSWER
18 percent. The gap wasn't about accuracy, it was about who knew they could ask for a review.
7 · THE REPLAY
Same overlapping-policy case, structured breakdown in place. What changes?
Tap to flip
ANSWER
The policy interaction gets flagged automatically instead of depending on Grace reconstructing it from memory. The correct, longer leave period gets approved, and the termination never happens.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this again for a different product. Which one, and who is the subject there?
Tap to flip
ANSWER
Corrigan Trust Bank's loan-modification assistant. The subject there is a struggling borrower who may never know a denial was ever contestable.
Check yourself Score: 0 / 0
Multiple choice
1. Why is a low overturn rate among new employees a warning sign, not good news?
A. New employees file more support tickets than veterans.
B. AskHR is specifically tuned to be more accurate for new hires.
C. It more likely reflects that new employees don't know a wrong answer can be challenged, not that answers for them are more correct.
D. New employees ask fewer questions overall.
Show hint
Look at the bar chart and the "unequal" step.
Show answer
C. The gap tracks awareness of the appeal process, not accuracy. A low number here hides risk instead of ruling it out.
True or false
2. True or false: asking AskHR about one policy clause at a time is always safer than asking about a whole case at once.
True
False
Show hint
Look at what happened the second time Grace used the clause-by-clause habit.
Show answer
False. Asking clause by clause can hide the interaction between policies, which is exactly what caused the missed leave extension the second time around.
Fill in the blank
3. Fill in the blank: the termination notice cited "___," not the actual policy-interaction mistake behind it.
Show hint
Look at the highlighted line in the story.
Show answer
Unauthorized absence. The real cause, a missed interaction between two leave policies, never appeared anywhere the worker could see it.
Short answer, where it wouldn't matter
4. Name an AskHR question where this level of governance concern is unnecessary.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: A routine question like how many vacation days remain this quarter, where a wrong answer is annoying but gets corrected the next time someone checks their balance.
Short answer, apply it yourself
5. Pick an internal company system you've encountered. Who would be least likely to know if it got something about them wrong?
Show hint
Think about someone newer, more junior, or less familiar with how to escalate a problem at that organization.
Show answer
Model answer: Often a new hire or contractor, unfamiliar with who to ask or what's normal, who has no baseline to notice that an automated answer looks off.
Short answer, name the reversal
6. What old decision does this answer take back, and why did it make sense when it was made?
Show hint
Look at "the decision I would take back."
Show answer
Model answer: Never building a structured way to review a multi-policy case. It made sense when most cases were simple, and broke the day a case genuinely wasn't.
Before you close the answer
Why this works
Tests whether you'll name a real power imbalance and a concrete fix, or default to a generic "add human review" answer that doesn't grapple with who actually can't push back.
Follow-up traps
"Isn't mandatory human sign-off just slower and more expensive?" Response: yes, on the subset of cases that touch pay or protected leave specifically, which is a deliberate trade of speed for the one place a wrong answer can't be walked back.
"Couldn't you just train employees on how to appeal?" Response: training helps, but it puts the burden on the exact group least equipped to notice something's wrong in the first place; the design has to carry more of that weight than the training does.
If pressed
Kestrel Mill's actual fix paired the mandatory sign-off with a structured, side-by-side clause view for any case touching more than one policy, so the interaction between rules is visible on one screen instead of depending on whichever reviewer happens to remember to cross-check by hand.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.