What does the EU AI Act require of a product like the one you last worked on?
- Route anything under the cutoff to a person before the rejection email goes out.Why: an automatic no with nobody looking is the whole problem. Every other fix sits downstream of this one.
- Tell candidates a machine read their file, and give them one real way to ask for a second look.Why: right now nobody outside the company can push back on the score at all, not because the score is wrong, but because nobody told them there was one.
- Log what matched and what was missing for every score, not just the ones somebody complains about.Why: without a plain record, nobody can go back a year later and explain why any one person was turned down.
- Test the training data for who it leaves out, before you switch the tool on.Why: it was built from the resumes of people who already got promoted, so it never learned what a career changer or a foreign employer's name looks like.
- Watch the callback rate by resume gap and by employer country, every month.Why: the overall pass rate looked fine the entire time. The whole problem lived inside one slice of it.
- Leave the hard, objective filters alone, like a legally required license.Why: not every automatic filter is a fairness risk, and human review is a limited resource better spent on the calls that actually involve judgment.
How to answer this, stage by stage
Eight moves. The first one says what the acronym actually means before any of the letters show up.
Let's learn
What happens when a computer decides you are not worth a callback, and it never has to say why?
Say a mid-size trucking company gets a tool that reads every resume for a job opening and scores it, so a recruiter only looks at the ones near the top.
Before the tool, three recruiters split six hundred resumes for a single warehouse supervisor posting. Each of them spent about fifteen hours reading, comparing notes, arguing over the close calls. Forty five hours of work, for one opening.
Now the tool reads all six hundred overnight. By morning it hands back a ranked list, and a recruiter spends under an hour on the top twenty. Everyone scored under fifty five gets a rejection email before a single person on the team opens their file.
Here is the part that matters. It is not that the tool makes mistakes. Every hiring process makes mistakes, a tired recruiter skips a good resume too. The real change is that nobody on the team ever sees the ones it turns down. For two years, one shape of candidate got filtered out overnight, and not one person checked a single file to see if that was fair.
At its worst, this ends up behind the slow way, not ahead of it. The old process was tiring and uneven, but a person read every resume once. This one can run for years getting the same kind of candidate wrong in the same way, and nobody finds out until a new hire asks an obvious question, or a regulator does.
What I would leave alone. Some filters should stay fully automatic. If a role legally needs a forklift certificate and a resume does not show one, reject it without a person in the loop. That is a fact check, not a judgment call, and there is no protected group hiding behind a yes or no license question. Save the human review for the calls that actually involve judgment.
The lesson. I used to think being compliant meant writing a policy that says the tool will not discriminate. It does not mean that. It means somebody can open any one rejected file, a year later, and reconstruct exactly why the machine said no. If you cannot do that, you do not have a hiring tool. You have a filing cabinet nobody can open.
Now here is the same thing as a story
The short version is above. Read this one when you want to feel why the fix matters, not just know what it is.
Marisol can read a resume's shape before she reads a single job title on it. Six years in the same seat will do that.
She joined Coastline Freight and Logistics as a recruiter and worked her way up to running the hiring team. For her first four years there, she read every resume that came in for a posting herself, all six hundred of them for something like the warehouse supervisor role, split across three people over a week.
Two years ago the company brought in the ranking tool. Marisol did not build it. Devon did, the recruiter who had her job before her, and by the time she inherited it, it was already just part of the desk. It cut three people's whole week down to under an hour. Good months followed, plenty of them. The team stopped arguing about close calls at five on a Friday.
For the first year, Marisol still pulled ten of the rejected resumes a month and read them herself, just to see. Nothing ever looked wrong. After about six months of that, she was down to three a quarter. By the start of her second year, she had stopped pulling any at all. The tool had never once looked wrong to her, and there is only so long you keep checking a thing that keeps being fine.
Then Owen started. Two weeks into the job, going through a shortlist, he asked her a question she did not have an answer for. Every single person on every shortlist for the last two quarters had come from one of about fifteen companies. Nobody from outside trucking and logistics. Nobody with a gap on their resume. Nobody who had worked abroad. He was not trying to catch her out. He had just noticed.
Marisol did not have an answer, so she pulled the real numbers instead of a sample. Every resume that scored under fifty five for the full quarter. Twelve hundred of them. She spent a weekend reading as many as she could by hand, and the pattern held on almost every page. A gap of six months or more. A last employer outside the country. A career change into logistics from somewhere else entirely. All of it scored low, again and again, no matter how strong the real experience underneath it was.
One of the twelve hundred belonged to a man named Solomon Getachew. He had spent six years running a shipping depot in Addis Ababa, twenty two people under him, no serious incident the whole time. Eight months before he applied, he had stepped back to care for his mother after a stroke. The tool read "Ethio Shipping and Logistics" and matched nothing it recognized. It read the gap and matched nothing there either. It scored him a forty one, and the rejection email went out that same afternoon. Solomon never knew a machine had read it at all. He assumed, the way people do, that something about him just was not quite right, and he quietly stopped listing that job on the next few applications he sent, in case it was the problem.
It was never really about the model getting one resume wrong. Getting one wrong is normal, a tired recruiter does that too. It was about a decision nobody in that room thought they were making twice: to let the score send its own answer, with nobody ever asked to double check the ones it turned away.
I was in the meeting where we made that call. It was not really a decision, it was a default nobody argued with. Under fifty five, the system sends the standard rejection, because reading every low score by hand seemed to defeat the entire point of building the tool. That was true. It was also the whole problem.
Run that same quarter again, but nothing under fifty five leaves on its own. All twelve hundred of those files sit in a queue Marisol opens herself, about two minutes a file. That is forty hours across a ten week quarter, roughly one working week, spread four hours at a time. Solomon's file is one of the forty she reads that Monday, instead of one of the eleven hundred and ninety nine nobody reads at all.
Same tool, same score. The only thing that changed is whose Monday the file lands on. I keep going back to that meeting, where reading every low score by hand felt like giving up on the whole point of the tool. I would argue the opposite this time. The point was never to remove the person from the process. It was to remove the six hundred resumes she did not need to read by hand, not the twelve hundred somebody should have.
Five letters, run against a resume instead of a case file
This is a risk question, so the framework is GUARD. A hiring score is a risk decision wearing a productivity feature's clothes, which is exactly why "we already have a person in the loop" gets said in meetings without anyone checking what that person actually sees.
And if you want to be sure it really works, try it somewhere else
A health insurer's tool scores every claim and flags some for extra review before it pays out. Different industry, same five letters, same trap.
G, groups. The claims adjuster who clears the review queue, and the policyholder whose claim just got held.
U, unequal. The holds land on claims with paperwork that does not look standard: an overseas provider, a translated invoice, a treatment code the model rarely sees. Routine domestic claims sail through.
A, ability to contest. There is a number to call about a denial. There is nothing to call about a hold, because a hold is not a denial. It just sits, and a mortgage payment does not wait for it.
R, reduce. During review, the tool can flag a claim, but it cannot hold it past five working days on its own. Past that, a person has to actively sign off on any more delay.
D, detect. Average days to payment, split by whether the claim names an overseas provider. Not the overall average. That one barely moves while one group waits three times as long.
Swap the trigger and it still runs
- Speed: hiring triples during a seasonal surge and nobody touches the cutoff score. Fifty five quietly filters out more people than it did in a normal month, because the same score now sits across a wider spread of resumes.
- Cost: the same tool, pointed at internal promotions instead of new hires. A wrong score no longer costs someone one job. It costs someone ten years of the career they were building inside the company.
- The model gets better: accuracy on the score climbs, so fewer files land in the review queue at all. The one person still checking it starts skimming, because it never finds anything anymore. Nothing about the design changed. The habit did.
Where people run it wrong
- Counting "a person is somewhere in the loop" as done, when that person only ever sees a monthly summary and never a single file.
- Testing overall accuracy and calling it bias testing, when the entire problem lives inside one slice the average never shows.
- Writing a fairness policy and treating the document itself as the fix, instead of changing what the software actually does on Monday morning.
If you're asked this cold
Ask what happens to the person who scores just under the cutoff, specifically, right now. Who opens their file, and how would they ever find out a machine touched it. That question buys you ten seconds, and it usually answers itself, because most teams never gave that person a name.
Flashcards (click a card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Compliance and legal partnership
- #2 Explain risk categorization under the EU AI Act in product terms.
- #3 How do you bring legal into an AI project early without slowing it down?
- #4 What questions will your legal team ask about training data, and how do you prepare?
- #5 Describe the copyright exposure of a generative feature.
- #6 How do you handle a customer contract that prohibits any use of their data for model improvement?
- #7 What disclosure obligations apply when users interact with an AI system?