ConceptAdvancedResponsible AI & Advanced Practice / Compliance and legal partnership / #1

What does the EU AI Act require of a product like the one you last worked on?

The direct answer
Do not let a score reject anyone by itself. Route anything under the cutoff to a person before the rejection goes out, tell candidates a machine read their file, and watch the callback rate for resumes with a gap or a foreign employer every month, not just the overall pass rate. That is what a hiring tool like this owes the EU AI Act: a person who can still say no to the machine, and a record good enough to show why it said no.
Do this, in order
  1. Route anything under the cutoff to a person before the rejection email goes out.Why: an automatic no with nobody looking is the whole problem. Every other fix sits downstream of this one.
  2. Tell candidates a machine read their file, and give them one real way to ask for a second look.Why: right now nobody outside the company can push back on the score at all, not because the score is wrong, but because nobody told them there was one.
  3. Log what matched and what was missing for every score, not just the ones somebody complains about.Why: without a plain record, nobody can go back a year later and explain why any one person was turned down.
  4. Test the training data for who it leaves out, before you switch the tool on.Why: it was built from the resumes of people who already got promoted, so it never learned what a career changer or a foreign employer's name looks like.
  5. Watch the callback rate by resume gap and by employer country, every month.Why: the overall pass rate looked fine the entire time. The whole problem lived inside one slice of it.
  6. Leave the hard, objective filters alone, like a legally required license.Why: not every automatic filter is a fairness risk, and human review is a limited resource better spent on the calls that actually involve judgment.

How to answer this, stage by stage

Eight moves. The first one says what the acronym actually means before any of the letters show up.

1
Say the plain version before the acronym shows up
Say it like this
"Under the EU AI Act, a tool that ranks job applicants counts as high risk, because it decides who gets a shot at a paycheck. High risk means five things have to be true: somebody checked who could get hurt before launch, a person can actually override it, there's a record good enough to explain any one decision, the training data got checked for who it leaves out, and the candidate can find out a machine touched their file and ask for another look."
Why this works
It gives the acronym real content in one breath, so you are never hiding behind the letters or the deadline dates.
2
Name the two people in the room, out loud
Say it like this
"There's the recruiter running the tool, and there's the applicant it's scoring. Only one of them has ever seen a dashboard."
Why this works
This is GUARD's G step, and it sets up exactly why the harm can sit unevenly later.
3
Say why this counts as high risk, not just software
Say it like this
"This isn't a spell checker. Hiring decides income. The Act puts sorting job applications in the same bucket as things like credit scoring, because getting it wrong follows someone for a year, not five minutes."
Why this works
It shows you understand the reasoning behind the label, not just the label itself.
4
Turn the score into a callback gap, not a percentage
Say it like this
"Resumes with a gap on them reached our shortlist score four percent of the time. Resumes without one reached it twenty two percent of the time. A score of forty one instead of sixty looks like grading. It's actually a callback, and under that, rent."
Why this works
It makes GUARD's U step concrete instead of abstract, and numbers land harder than a claim of unfairness.
5
Say what a person can actually do about it
Say it like this
"Right now, anyone under the cutoff gets one email and that's the whole interaction. No score shown, no reason given, no way to ask a person to look again. That's the real gap here, not the model's accuracy."
Why this works
This is GUARD's hardest and strongest step. Naming who can't push back is what separates a risk answer from a rollout answer.
6
Name the one thing you'd build, not the policy you'd write
Say it like this
"I'd stop the score from sending its own rejection. Anything under the cutoff goes into a queue a person opens first, and the email itself says what matched and what didn't. 'Matched: five years supervisory experience. Missing: a cold chain certificate.' Not silence."
Why this works
A design a person can build on Monday beats a policy nobody can point to a year later.
7
Say how you'd catch it without waiting for a complaint
Say it like this
"Every month I'd pull the pass through rate split by whether a resume shows a gap, and by whether the last employer is outside the country. Not the overall rate, that one barely moves. And once a quarter I'd have someone re read fifty of the auto rejected files blind, no score shown, and see how many they'd have kept."
Why this works
This is the detect step, and it's the only thing that catches drift before a candidate, a journalist, or a regulator does.
8
Close on the sentence that isn't hedging
Say it like this
"So: don't let a score reject anyone by itself, tell candidates a machine touched their file and give them a real way to ask again, and watch the callback gap every month instead of trusting the overall number."
Why this works
Restating the decision in one breath is the part the interviewer actually remembers.

Let's learn

What happens when a computer decides you are not worth a callback, and it never has to say why?

Say a mid-size trucking company gets a tool that reads every resume for a job opening and scores it, so a recruiter only looks at the ones near the top.

Knowledge spark: what makes an AI system "high risk" Not every AI tool gets extra rules under the EU AI Act. The law saves that for systems that can change something big in someone's life: who gets hired, who gets a loan, who gets flagged at a border. A tool that ranks job applicants sits right in that group, so it has to clear a higher bar before anyone switches it on.

Before the tool, three recruiters split six hundred resumes for a single warehouse supervisor posting. Each of them spent about fifteen hours reading, comparing notes, arguing over the close calls. Forty five hours of work, for one opening.

Now the tool reads all six hundred overnight. By morning it hands back a ranked list, and a recruiter spends under an hour on the top twenty. Everyone scored under fifty five gets a rejection email before a single person on the team opens their file.

Here is the part that matters. It is not that the tool makes mistakes. Every hiring process makes mistakes, a tired recruiter skips a good resume too. The real change is that nobody on the team ever sees the ones it turns down. For two years, one shape of candidate got filtered out overnight, and not one person checked a single file to see if that was fair.

At its worst, this ends up behind the slow way, not ahead of it. The old process was tiring and uneven, but a person read every resume once. This one can run for years getting the same kind of candidate wrong in the same way, and nobody finds out until a new hire asks an obvious question, or a regulator does.

One resume, and what never happened to it: never opened by a person, no reason sent back, no way to ask again
One of twelve hundred files like this a quarter
The decision I would take back We let any score under fifty five send its own rejection, straight to the candidate, no queue, no person. It made sense at the time. Reading every low score by hand felt like giving up on the entire reason we built the tool. It also meant nobody ever checked the ones it turned away, for two full years.

What I would leave alone. Some filters should stay fully automatic. If a role legally needs a forklift certificate and a resume does not show one, reject it without a person in the loop. That is a fact check, not a judgment call, and there is no protected group hiding behind a yes or no license question. Save the human review for the calls that actually involve judgment.

Solomon's file was never wrong. It was never read.

The lesson. I used to think being compliant meant writing a policy that says the tool will not discriminate. It does not mean that. It means somebody can open any one rejected file, a year later, and reconstruct exactly why the machine said no. If you cannot do that, you do not have a hiring tool. You have a filing cabinet nobody can open.

Now here is the same thing as a story

The short version is above. Read this one when you want to feel why the fix matters, not just know what it is.

Marisol can read a resume's shape before she reads a single job title on it. Six years in the same seat will do that.

She joined Coastline Freight and Logistics as a recruiter and worked her way up to running the hiring team. For her first four years there, she read every resume that came in for a posting herself, all six hundred of them for something like the warehouse supervisor role, split across three people over a week.

Two years ago the company brought in the ranking tool. Marisol did not build it. Devon did, the recruiter who had her job before her, and by the time she inherited it, it was already just part of the desk. It cut three people's whole week down to under an hour. Good months followed, plenty of them. The team stopped arguing about close calls at five on a Friday.

For the first year, Marisol still pulled ten of the rejected resumes a month and read them herself, just to see. Nothing ever looked wrong. After about six months of that, she was down to three a quarter. By the start of her second year, she had stopped pulling any at all. The tool had never once looked wrong to her, and there is only so long you keep checking a thing that keeps being fine.

Then Owen started. Two weeks into the job, going through a shortlist, he asked her a question she did not have an answer for. Every single person on every shortlist for the last two quarters had come from one of about fifteen companies. Nobody from outside trucking and logistics. Nobody with a gap on their resume. Nobody who had worked abroad. He was not trying to catch her out. He had just noticed.

Marisol did not have an answer, so she pulled the real numbers instead of a sample. Every resume that scored under fifty five for the full quarter. Twelve hundred of them. She spent a weekend reading as many as she could by hand, and the pattern held on almost every page. A gap of six months or more. A last employer outside the country. A career change into logistics from somewhere else entirely. All of it scored low, again and again, no matter how strong the real experience underneath it was.

Marisol did not stop checking the file. She stopped being handed one to check.

One of the twelve hundred belonged to a man named Solomon Getachew. He had spent six years running a shipping depot in Addis Ababa, twenty two people under him, no serious incident the whole time. Eight months before he applied, he had stepped back to care for his mother after a stroke. The tool read "Ethio Shipping and Logistics" and matched nothing it recognized. It read the gap and matched nothing there either. It scored him a forty one, and the rejection email went out that same afternoon. Solomon never knew a machine had read it at all. He assumed, the way people do, that something about him just was not quite right, and he quietly stopped listing that job on the next few applications he sent, in case it was the problem.

Score under fifty five, now, versus score under fifty five with a queue
The pause we took out

It was never really about the model getting one resume wrong. Getting one wrong is normal, a tired recruiter does that too. It was about a decision nobody in that room thought they were making twice: to let the score send its own answer, with nobody ever asked to double check the ones it turned away.

I was in the meeting where we made that call. It was not really a decision, it was a default nobody argued with. Under fifty five, the system sends the standard rejection, because reading every low score by hand seemed to defeat the entire point of building the tool. That was true. It was also the whole problem.

Run that same quarter again, but nothing under fifty five leaves on its own. All twelve hundred of those files sit in a queue Marisol opens herself, about two minutes a file. That is forty hours across a ten week quarter, roughly one working week, spread four hours at a time. Solomon's file is one of the forty she reads that Monday, instead of one of the eleven hundred and ninety nine nobody reads at all.

Same tool, same score. The only thing that changed is whose Monday the file lands on. I keep going back to that meeting, where reading every low score by hand felt like giving up on the whole point of the tool. I would argue the opposite this time. The point was never to remove the person from the process. It was to remove the six hundred resumes she did not need to read by hand, not the twelve hundred somebody should have.

Five letters, run against a resume instead of a case file

This is a risk question, so the framework is GUARD. A hiring score is a risk decision wearing a productivity feature's clothes, which is exactly why "we already have a person in the loop" gets said in meetings without anyone checking what that person actually sees.

GUARD: groups, unequal, ability to contest, reduce, detect
GUARD, for risk, safety and fairness questions
G, groups. Two people, and the law wants both named. The operator is Marisol, who runs the tool and can see every score it gives. The subject is Solomon, who the tool scores and who has never seen a dashboard in his life.
U, unequal. The resumes under fifty five are not a random slice of the applicant pool. They cluster hard around one shape: a gap, a foreign employer, a career change. Split the same pool that way and the gap shows up in the numbers, not just the stories.
Reaching the shortlist score, by resume shape
Same quarter, same applicant pool. A score of fifty five or higher is needed to reach the shortlist.
Gap of six months or more
4%
No gap on the resume
22%
Nearly six times the reach rate, and the overall pass rate never showed it. It stayed level the entire two years, because it was always an average of a group that mostly cleared the bar and a group that almost never did.
A, ability to contest. Solomon cannot appeal a decision, because as far as he knows there was not one. Nobody told him a machine read his file, so there is nothing to push back on. The rejection email reads the same whether a person or the tool sent it, on purpose, so nobody would feel singled out. It also means nobody ever finds out they were.
R, reduce. Anything under the cutoff goes into a queue. A person opens the file before any rejection goes out, and the email itself says what matched and what did not.
D, detect. The pass through rate, split by resume gap and by employer country, checked every month. Not the overall rate. That one stayed normal for two straight years.
Where this answer would fail If the fix here were a fairness policy, a training deck, or a review board, none of it counts. "A person opens the file before it rejects" is a build ticket. Somebody can ship it Monday, and you can check whether they did.

And if you want to be sure it really works, try it somewhere else

A health insurer's tool scores every claim and flags some for extra review before it pays out. Different industry, same five letters, same trap.

G, groups. The claims adjuster who clears the review queue, and the policyholder whose claim just got held.
U, unequal. The holds land on claims with paperwork that does not look standard: an overseas provider, a translated invoice, a treatment code the model rarely sees. Routine domestic claims sail through.
A, ability to contest. There is a number to call about a denial. There is nothing to call about a hold, because a hold is not a denial. It just sits, and a mortgage payment does not wait for it.
R, reduce. During review, the tool can flag a claim, but it cannot hold it past five working days on its own. Past that, a person has to actively sign off on any more delay.
D, detect. Average days to payment, split by whether the claim names an overseas provider. Not the overall average. That one barely moves while one group waits three times as long.

Swap the trigger and it still runs

  • Speed: hiring triples during a seasonal surge and nobody touches the cutoff score. Fifty five quietly filters out more people than it did in a normal month, because the same score now sits across a wider spread of resumes.
  • Cost: the same tool, pointed at internal promotions instead of new hires. A wrong score no longer costs someone one job. It costs someone ten years of the career they were building inside the company.
  • The model gets better: accuracy on the score climbs, so fewer files land in the review queue at all. The one person still checking it starts skimming, because it never finds anything anymore. Nothing about the design changed. The habit did.

Where people run it wrong

  • Counting "a person is somewhere in the loop" as done, when that person only ever sees a monthly summary and never a single file.
  • Testing overall accuracy and calling it bias testing, when the entire problem lives inside one slice the average never shows.
  • Writing a fairness policy and treating the document itself as the fix, instead of changing what the software actually does on Monday morning.

If you're asked this cold

Ask what happens to the person who scores just under the cutoff, specifically, right now. Who opens their file, and how would they ever find out a machine touched it. That question buys you ten seconds, and it usually answers itself, because most teams never gave that person a name.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
Which framework fits a question about what a law requires of an AI product, and why?
Tap to flip
ANSWER
GUARD, for risk, safety and fairness. The Act is really asking who a hiring score can hurt and who never gets to push back on it, which is exactly what GUARD is built to find.
2 · THE TWO PEOPLE
Name the two people this answer names, and which one can see the product.
Tap to flip
ANSWER
Marisol Reyes, the recruiter who runs the tool, and Solomon Getachew, an applicant it scored a forty one. Only Marisol has ever opened a dashboard.
3 · THE HABIT
What did Marisol quietly stop doing, and over how long?
Tap to flip
ANSWER
Pulling ten rejected resumes a month to read herself. Ten a month became three a quarter became none, over about eighteen months, because the tool never once looked wrong.
4 · WHAT THE LAW ASKS FOR
Name the concrete things a hiring tool like this has to have, in plain words.
Tap to flip
ANSWER
A person who can still say no to it. A record good enough to explain any one decision later. The training data checked for who it leaves out. And a candidate who can find out a machine touched their file and ask for another look.
5 · THE OLD DECISION
What old decision does this answer take back, and why did it make sense at the time?
Tap to flip
ANSWER
Letting any score under fifty five send its own rejection, no queue, no person. It made sense because reading every low score by hand seemed to defeat the point of building the tool. It also meant nobody ever checked the ones it turned away.
6 · THE NUMBER
Fill in: resumes with a gap reached the shortlist score of fifty five or higher ______ percent of the time, against ______ percent for resumes with no gap.
Tap to flip
ANSWER
Four percent, against twenty two percent. The overall pass rate never showed this. It was buried inside one slice of the applicants.
7 · THE REPLAY
Same quarter, new design. What changes, and by how much?
Tap to flip
ANSWER
Nothing under fifty five leaves on its own. Marisol opens all twelve hundred queued files herself, about two minutes each, forty hours across the quarter. Solomon's file is one of the ones she actually reads.
8 · TRANSFER
Section four runs GUARD again on a different product. Which one, and what does the reduce step become?
Tap to flip
ANSWER
A health insurer's claims tool. Reduce: a claim can be flagged for extra review, but nobody can hold it past five working days without a person actively signing off on the delay.

Check yourself Score: 0 / 0

Multiple choice
1. A recruiter tells you the resume tool auto rejects anything under a score of fifty five. Per the reduce step, what's the first change you'd make?
  • A. Lower the cutoff to forty five.
  • B. Route anything under the cutoff to a person before any rejection goes out.
  • C. Retrain the model on more resumes and check back next quarter.
  • D. Show candidates their exact numeric score.
Show hint
The reduce step is a specific design change, not a bigger or smaller version of the same automatic step.
Show answer
B. A and C both just adjust a dial on the same automatic system. D sounds transparent, but a bare number with no reason still gives nobody a real way to push back. Only routing it through a person changes what actually happens to the file.
True or false
2. True or false: because the tool's overall pass rate looked normal the whole time, there wasn't really a bias problem.
  • True
  • False
Show hint
Where would a problem in one small slice of applicants show up in an average of everyone?
Show answer
False. The overall rate can look completely healthy while one group, resumes with a gap or a foreign employer, gets a wildly different outcome underneath it. An average is exactly the place a problem like this hides.
Fill in the blank
3. Resumes with a gap of six months or more reached the shortlist score of fifty five or higher ______ percent of the time, against ______ percent for resumes with no gap.
Show hint
It's the pair of numbers behind the two-bar chart in the GUARD section.
Show answer
Four percent, against twenty two percent. That gap sat inside the data for two years, and nobody saw it because nobody was looking at anything other than the overall number.
Short answer
4. Name a filter in this same hiring pipeline that you would leave fully automatic, and say why that one is safe.
Show hint
Look for a filter that checks an objective fact instead of making a judgment call.
Show answer
Model answer: "A legally required license or certification, like a forklift certificate for a role that needs one by law. It's a yes or no fact, not a judgment about fit, so there's no protected group hiding behind it and no reason to spend a person's limited review time checking it. Save the human review for the calls that involve judgment, like an unfamiliar employer or a gap."
Multiple choice
5. A team says: "We already have a person in the loop, they check the flagged resumes once a month in a summary report." What's wrong with calling that meaningful human oversight?
  • A. Nothing, monthly checks are frequent enough for any hiring tool.
  • B. The person never sees an actual file, only a summary, so there's nothing for them to actually override.
  • C. The report should be weekly instead of monthly.
  • D. The model needs more training data before this even matters.
Show hint
Ask what the person in the loop can actually change, not how often they look.
Show answer
B. A summary report lets someone notice a pattern eventually, but it gives them no lever on any single file before the rejection already went out. Real oversight means a person can stop or change one specific decision, not just read about a trend after the fact.
Short answer, apply it yourself
6. Pick an AI system you've used or built. Who's the person it affects who never sees it and has no way to push back?
Show hint
Look for someone downstream of the output who isn't the one operating the tool and can't opt out.
Show answer
Model answer: "A chatbot that triages support tickets by urgency at a company I worked with. The agent sees the model's tag. The customer never learns their ticket got marked low priority by a script, and there's no way for them to ask for a human read instead, they just wait longer with no idea why." Any answer works if you can name someone who absorbs the tool's mistake and has no lever to pull.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more