CaseIntermediateAI Opportunity & Model Strategy / Opportunity identification for AI / #4

Rank these by AI suitability and justify: expense approval, contract review, invoice matching, hiring decisions.

[LEAD] · why one confidence rule can't grade four different jobs, tested on Elderbrook's travel platform Driftwick

Driftwick is Elderbrook's travel and expense platform. It books flights and hotels, matches travel invoices to the trip they belong to, and drafts expense reports so nobody has to open a spreadsheet by hand. Ilonka Vantresca owns the small AI team that builds all of that. This quarter, four different requests landed on her roadmap at once: keep matching invoices, start approving expense reports, start flagging risky lines in hotel contracts, and, from Kalju Norrbin, Elderbrook's head of people, help decide who gets hired. Only one of those four was ever going to be safe to hand over completely.

The direct answer
Rank the four by what a wrong answer actually costs and how fast anyone would catch it, not by how excited the person asking happens to be. In that order: invoice matching first, expense approval second, contract review third, hiring decisions last. Build in that order, and even the task at the top still keeps a person checking anything past a real dollar line, because ranking high buys more freedom. It never buys zero oversight, anywhere on the list, especially the bottom.
Do this, in order
  1. Rank invoice matching, expense approval, contract review, then hiring decisions, in that order.Why: this is the direct answer, just laid out as a list.
  2. Let invoice matching run itself under a real dollar line.Why: a wrong one is cheap to fix and gets caught the same day, so it earns the most freedom.
  3. Keep expense approval assisted: the model approves, a manager still sees anything unusual.Why: real money and trust sit on this one, and policy rules alone miss the odd real case.
  4. Let contract review flag. Never let it sign.Why: one missed clause can cost hundreds of thousands, and it might not get noticed for six months.
  5. Keep hiring fully human. The model can point out a missing skill. It never decides.Why: the person harmed by a wrong call can't push back, and the mistake can hide for a whole quarter.
  6. Kill the one confidence rule running every model call the same way. Give each task its own.Why: ninety percent sure means something different on a twelve dollar invoice than it does on somebody's job.

How to answer this, stage by stage

Nobody is grading whether you can name four tasks in order. They're grading whether you can say, out loud, what actually earns a task a spot on this list.

1
Ground it in one team's real request queue
Say it like this
"Let's ground this. Elderbrook builds Driftwick, a travel and expense platform. The AI team, that's Ilonka Vantresca and two engineers, gets asked to automate four different jobs this quarter: matching invoices, approving expense reports, reviewing hotel contracts, and, from the head of people, deciding who moves forward in hiring. I'll rank those four, not talk about AI suitability as an idea floating in the air."
Why this works
Naming a real request queue stops the ranking from turning into a list of abstractions nobody could act on.
2
State the method in one breath
Say it like this
"I'll run this as LEAD. Link ties 'AI suitable' to something real: how bad a wrong answer is, how fast you'd catch it, and who gets hurt. Early signal is where the actual ranking comes from. Abuse is how people twist a good ranking into 'ship it everywhere with no review.' Decision is what I actually do once I have the order."
Why this works
Two seconds of structure tells the interviewer this is a method, not a gut-feel list dressed up after the fact.
3
Give the order before any reasoning
Say it like this
"Short version, in order: invoice matching first, expense approval second, contract review third, hiring decisions dead last. If I stopped talking right now, that's the answer."
Why this works
The interviewer never has to dig through the reasoning to find the actual ranking.
4
Name what "AI suitable" actually measures
Say it like this
"It isn't accuracy. Every one of these four models can hit ninety five percent in a demo. What changes is three things: how bad one wrong answer is, how fast and cheap it is to catch, and who gets hurt if nobody catches it in time."
Why this works
This is the reframe. Without it, the ranking looks like a guess wearing a fancy word.
5
Walk the four tasks against those measures
Say it like this
"Invoice matching: thirty eight thousand a month, a wrong one costs about twelve dollars to fix, and finance catches it the same day, so it can run with almost nobody watching. Expense approval: a wrong auto approval costs a few hundred dollars, caught inside about a week once a manager reviews it, so it needs a person on anything unusual. Contract review: one missed clause in a hotel deal can cost hundreds of thousands, and it might sit there for six months before anyone notices, so the model only gets to flag, never sign. Hiring: a wrong reject can end someone's shot at a job, it's almost impossible to undo, and the only thing checking for it is an audit that runs once a quarter. So a model doesn't get a vote there at all."
Why this works
This is the real content of the early signal step, the actual comparison the interviewer is testing for.
6
Catch the way this ranking gets twisted
Say it like this
"Here's the trap. Someone hears 'invoice matching is AI safe' and hears 'no review anywhere is fine now, because the model's good.' That's backwards. High on this list means safe to trust with more freedom. It never means safe to trust with none, anywhere on the list, especially at the bottom, where the harm is heaviest."
Why this works
Naming the misuse out loud is what separates a candidate who understands the ranking from one who just memorized an order.
7
Close on the sequencing decision, with the boundary named
Say it like this
"So here's what I actually do. I build invoice matching first, because it proves value fastest and it's the safest place to let the model run alone, above a threshold. Then expense approval, then contract flagging. Hiring stays helper only, not just for this quarter, for good. And even on invoice matching, anything past a real dollar amount still goes to a person, because ranking first isn't the same as being risk free."
Why this works
This is the actual decision, said once, plainly, with the line drawn in the right place.

Let's learn

What happens when four different teams ask the same small AI group to automate four completely different jobs, all at once?

Driftwick is Elderbrook's travel and expense platform. Book a flight, book a hotel, and Driftwick's AI matches the invoice to the trip it belongs to and drafts the expense report before anyone has to open a spreadsheet.

Hand sketched labeled parts diagram titled One team, four very different jobs, with a central box icon labeled Driftwick AI Team, and four labeled callouts radiating outward: match travel invoices, approve expense reports, review hotel contracts, and decide who gets hired.
One small team. Four requests that look similar on a roadmap slide and are nothing alike underneath.

When Driftwick's AI first went live, invoice matching was the only thing running. Ilonka's team wrote one rule for it: if the model was more than ninety percent sure, and the amount was small enough that a mistake wouldn't matter, let it act alone. Check everything else once a month. That rule automatically matched 38,000 invoices a month, and it saved the finance team about three hundred hours of manual matching every month, real hours, given back.

Over the next year, three more requests landed on that same rule without anyone sitting down to ask whether it still fit. Expense approval got wired to it in spring. A first pass at flagging risky contract clauses got wired to it in summer. And, quietly, in September, Kalju Norrbin's hiring pipeline got an auto reject step built on the exact same rule too, to clear a three week backlog of applications before a hiring freeze made the delay somebody else's problem.

We did not build a biased hiring tool on purpose. We built one confidence rule, once, and let it decide four completely different things it was never actually tested against.

None of the four broke because the model got worse. That's not the story here. Ninety percent sure, act alone, means something different depending on what sits behind it. On an invoice it means fixing a twelve dollar typo alone. On a hiring decision it means ending someone's application alone, and not finding out for months if it was even the right call.

Hand sketched numbered icon list titled Ranked, how much can a model be trusted here. Item one, invoice matching, model can act on its own, green icon. Item two, expense approval, model acts, a person checks, amber icon. Item three, contract review, model flags, never signs, amber icon. Item four, hiring, model helps, a person decides, red icon.
Same four jobs. Sorted by what a wrong answer actually costs, not by who asked first.
All four tasks, plotted on stakes versus how long a wrong answer hides
1 day 10 days 100 days Days before a wrong answer gets caught $100 $1,000 $10,000 $100,000 Cost if the answer is wrong Invoice matching Expense approval Contract review Hiring decisions
Invoice matching Expense approval Contract review Hiring decisions
Both axes use a log scale, so equal steps mean ten times more, not the same amount more. Bottom left is the safe corner: cheap to be wrong, fast to notice. Top right is where invoice matching's rule should never have gone near.

Line up the real numbers behind those four dots. Invoice matching: about 38,000 a month, a wrong one costs about twelve dollars to fix, and finance sees it the same business day, because the match either lines up against the booking record or it doesn't. Expense approval: a wrong auto approval costs a few hundred dollars, and a weekly manager review catches it inside about four days. Contract review: Elderbrook signs or renews about sixty hotel and vendor rate deals a year, each worth around $420,000 in yearly spend, and a missed clause can sit unnoticed for about six months, roughly 180 days, until a hotel invokes it or a renewal date quietly passes. Last year, a missed auto renewal clause on one hotel block locked Elderbrook into an old rate for another twelve months once the window to renegotiate had already closed. Real cost: about $180,000.

Hand sketched comparison diagram titled Which mistake gets caught first. Left panel, a gauge icon labeled Invoice matching, caption a bad match is caught the same day. Right panel, a question mark box icon labeled Hiring decisions, caption a bad reject can hide for months.
One of these mistakes gets caught before lunch. The other one waits for a scheduled check that runs four times a year.

Hiring sits furthest from safe, and here is why. The check that would catch a bad call, Elderbrook's routine equal opportunity audit, only runs once a quarter, about ninety days between checks. This time it happened to land at eleven weeks, seventy seven days, a bit of luck in the calendar, not something the design earned. The audit found that candidates with a gap in their work history were being rejected at nearly double the rate of otherwise similar candidates, sixty one percent against thirty four percent. Employment gaps track with caregiving leave. Caregiving leave tracks, more than most places want to admit, with gender.

Hand sketched timeline titled Eleven weeks, one routine check, with four milestones: Pilot goes live, same rule as invoices. Reject pattern starts, nobody is watching yet. Weeks pass, no alert exists for this. Audit finds it, week eleven, by luck, this last milestone shown in red as emphasized.
Nobody decided to look away. Nothing about the design gave anyone a reason to look sooner.

About 640 applications went through that stage in eleven weeks. Around sixty of the auto rejected ones should have gone to a person first. Untangling it, the legal review, reopening the cases, rebuilding the process, ran to about $340,000, and that is before counting what it costs a company to sit under an equal opportunity finding.

The choice I would take back One confidence rule, ninety percent sure, act alone, check the rest once a month, used for every model call on the whole platform. It made sense the day it was written, because invoice matching was the only thing live, and one rule was simpler to build than four. It stopped making sense the day a second, far higher stakes task got wired to the same switch, and nobody said so out loud.

What I would leave alone: Driftwick's AI also decides which hotels and flights show up first in a traveler's search results, with zero review, and that is exactly right. A ranked list is a suggestion. Nobody is harmed by seeing the third best hotel before the best one, they just scroll past it. Being wrong there costs nothing, because nothing was ever decided, only offered.

The lesson: "AI suitable" is not one number set once for a whole platform. It is a separate call for every task, about how bad a wrong answer is and how fast anyone would know. Treat every model call as the same call, and a genuinely useful invoice tool ends up quietly deciding something about somebody's job that nobody meant it to decide.

Now here is the same thing as a story

Use this version when you have the time. The short version above is what you would actually say out loud in the room. This one is for feeling why the order matters.

Ilonka can tell which AI request is actually simple, and which one is a lawsuit with a nice screen on top, usually before the person asking finishes their sentence. She has run Driftwick's AI team for three years, and she built the very first version of the platform's one automated rule herself: read a hotel or airline invoice, match it against the trip it belongs to, let the match through alone if the model is more than ninety percent sure and the amount is small enough that a mistake would not matter. It was the only automated thing on the whole platform, and it worked so well that finance stopped opening the matching queue most mornings.

For a year, that felt like the whole job. New requests kept landing, and each one got quietly pointed at the same rule, because the rule already existed and it already worked. Nobody sat down and asked whether ninety percent sure, act alone, should mean the same thing for a fifty dollar taxi receipt as it does for a four hundred thousand dollar hotel contract. It was just the rule. It was right there.

Then, one afternoon in September, Kalju Norrbin came by with a backlog problem. Elderbrook's own early stage resume screen was three weeks behind, and a hiring freeze was about to make that delay somebody else's fault instead of his. Ilonka's team had the rule sitting right there, already tested, already running three other jobs. It took an afternoon to point it at resumes instead of receipts. Nobody meant any harm by it. It was the fastest fix in the building.

For eleven weeks it ran quietly, doing exactly what it was told. Nobody watched it the way finance watched the invoice queue, because on the screen, nothing about it looked different from the other three jobs already running the same rule.

Then the quarterly audit came around, the way it always does, checking a random sample of hiring decisions for basic fairness. This time the sample flagged something real: candidates with a gap in their work history were being auto rejected at nearly double the rate of otherwise similar candidates. Nobody had done anything wrong on purpose. The model had learned, from years of past hiring data, that gaps tracked weakly with worse outcomes down the line, and ninety percent sure turned that weak pattern into reject, no review, every time.

A wrong invoice costs twelve dollars and gets caught by lunch. A wrong reject costs somebody a job they never knew they lost, and it took a scheduled audit to notice, not because nobody cared, but because nothing about the design gave anyone a reason to look sooner.

Kalju, to his credit, did not try to talk his way around the number once it was in front of him. His question was the honest one: "If the model's good enough to run invoice matching alone, why can't it run this alone too, once we fix the gap issue?" Ilonka told him no, and had to say why more than once. Being good enough to run alone on an invoice was never really about how smart the model was. It was about a wrong invoice costing twelve dollars and getting caught the same afternoon. A wrong reject cost somebody a chance they never knew they'd lost, and the only thing checking for it ran four times a year.

Run the same eleven weeks again, with the new rule instead of the old one. The model still flags the same six hundred forty resumes, for the same reasons. But every reject now goes to a person before anyone's application actually closes. The worst case stops being sixty people rejected wrongly and found three months late. It becomes zero, because the model was never holding the pen. It was only pointing at the page.

What I'd tell myself, back when that first rule felt so clean it seemed almost wasteful not to reuse it: the rule was never wrong. It was right, for invoices. Reusing it everywhere was not being efficient. It was assuming four different jobs were the same job, because they all happened to run through the same kind of model.

LEAD, so "AI suitable" means something you can check, not a feeling

Not a way to prove a model is smart enough. LEAD forces you to say exactly what a wrong answer costs, and exactly how fast you would know, before anything gets to run alone.

LLink. What "AI suitable" actually has to connect to.
Not a demo accuracy number, and not how excited the person asking happens to be. Three real things: how bad one wrong answer is, how fast and cheap it is to catch, and who gets hurt if nobody catches it in time.
Every one of Driftwick's four candidate jobs can hit ninety five percent accuracy in a demo. What actually decides whether that is safe is how bad the other five percent is, and how fast anyone would know it happened.
EEarly signal. Where the actual order comes from.
Line them up, cheapest and fastest first. Invoice matching: about 38,000 a month, twelve dollars to fix a wrong one, caught the same business day. Expense approval: a few hundred dollars, caught in about four days. Contract review: hundreds of thousands on a $420,000 contract, up to about 180 days before anyone notices. Hiring: the only check that would catch a bad call runs once a quarter, about ninety days, and the harm has already landed on a real person by the time it does.
The invoice matching rule caught its own mistakes by lunch. The hiring rule, wired to the exact same math, waited eleven weeks for a scheduled check to notice.
Hand sketched comparison diagram titled Which mistake gets caught first, shown again. Left panel, invoice matching, caught the same day. Right panel, hiring decisions, can hide for months.
The same picture again, on purpose. It's the whole reason invoice matching and hiring can't share one rule.
Days before a wrong answer gets caught, by task
50 100 150 0 0.5 Invoice matching 4 Expense approval 180 Contract review 90 Hiring decisions
InvoiceExpenseContractHiring
Invoice matching's bar is basically invisible on this scale. That is the whole point: it gets caught before anyone would even call it a delay.
AAbuse. How this ranking gets misused.
The trap: someone hears "invoice matching earns full freedom" and hears "this whole ranking is a green light to skip review, starting from the top." It is backwards. High on the list means the task can survive being trusted with more freedom. It never means every task on the list gets the same amount of it, especially the one at the bottom, where the harm is heaviest and hardest to undo. There is a second, quieter way it gets gamed too: build the hiring eval only from clearly qualified and clearly unqualified resumes, where the call is obvious, and it will score just as clean as invoice matching on paper. The real risk was never in the obvious cases. It lived in the ambiguous ones, exactly where a real audit, not a tidy eval set, actually found it.
Kalju's version of the trap, said out loud: "if invoice matching runs alone, hiring should too, once it's good enough." Good enough was never the question.
Hand sketched comparison diagram titled Same ranking, one wrong reading. Left panel, a scale icon labeled What the ranking says, caption trust invoice matching with more freedom. Right panel, a question mark box icon labeled What got asked for, caption let hiring run with no one watching.
Two readings of the exact same list. Only one of them a person on the receiving end would survive.
DDecision. What actually happens to the roadmap.
Build in the ranked order: invoice matching stays live and gets tuned, expense approval next, contract flagging after that. Hiring gets rebuilt as helper only, not paused, not killed, just never handed the final call. And the same sentence goes to everyone asking for a piece of the roadmap, including the one whose task ranked first: ranking high buys more freedom. It never buys all of it.
Above a real dollar line, invoice matching still goes to a person too. Being first on the list is not the same as being risk free.
Hand sketched decision tree titled What happens after the model answers. Root: a model finishes its answer. Four branches: an invoice with a small amount and high confidence leads to the model acting alone. An expense report with anything unusual leads to a manager checking it. A clause found in a contract leads to it being flagged for a lawyer. A hiring signal leads to it being shown to a person, who decides.
Same model, same platform. Four different amounts of freedom, and only one of them lets it act without asking first.

The recap, one line per letter: link "AI suitable" to how bad a wrong answer is, how fast anyone would catch it, and who gets hurt, not to a demo accuracy score. The early signal is the actual comparison across the four tasks, and it is why invoice matching leads and hiring trails, by a wide margin. Name the abuse plainly: treating a high rank as permission to skip review everywhere, and building an eval set that only ever shows the easy cases. And the decision is what makes any of this real: build in the ranked order, and keep a person in the loop at the top of the list too, not just the bottom.

Two things worth naming outright, since the real judgment sits here. Kalju's ask, to let the hiring model decide once it "got good enough," was taken seriously and turned down, not because the team distrusted the model's numbers, but because a rejected candidate has no way to ever push back on a call they never even knew got made, and a scheduled quarterly audit is the fastest anyone would notice something went wrong. The AI specific failure worth naming by name is exactly what the audit caught: a fairness gap that never shows up in an aggregate accuracy number, only in a check broken out by group, run on a schedule, not waited on. And the trade off was real and named on purpose: Ilonka could raise invoice matching's bar from ninety to ninety eight percent sure, making the auto run rule even safer, but that pushes thousands more invoices into the manual queue every month, trading a rare twelve dollar mistake for a very real chunk of finance's week. Ninety percent, under fifty dollars, stayed the line, because on this specific task, that is where the math actually holds.

And if you want to be sure it really works, try it somewhere else

Same four letters, a waiting room instead of a departures board, and this time the highest stakes task is about an animal that cannot say what is wrong.

Yewgarden Veterinary Group runs eleven clinics, and its AI layer, Mossglen, handles four requests of its own. Rafailia Ulfsson, the group's operations director, got asked to rank the same way Ilonka did: appointment reminders sent automatically, visit note summaries turned from a vet's spoken record into a structured chart, insurance claims auto filled and submitted, and phone in triage, deciding which symptoms need to be seen today and which can wait for a scheduled visit.

Reminders rank first: about 12,000 sent a month, next to no cost if one is wrong, since a client just calls to reschedule, and the mistake is obvious within the hour. Visit summaries come next: about 3,000 visits a month, a wrong summary might mean a missed detail costing around $150 in a repeat exam, caught within about a week when the vet reviews the chart before a follow up. Insurance claims sit third: about 800 a month, a wrong submission risks a denial or an overpayment worth about $220, caught in three to four weeks once the insurer processes it. Triage sits last, by a wide margin: the stakes are an animal's health, the model cannot ask a follow up question the way a vet tech would, and a wrong "it can wait" call might not get noticed for weeks, if it ever gets traced back to the call at all. The pet owner it fails has no way to contest a decision made over the phone in real time.

Hand sketched numbered icon list titled Same method, a vet clinic's four jobs. Item one, appointment reminders, runs alone, easy to check, green icon. Item two, visit summaries, vet skims it, fixes gaps, amber icon. Item three, insurance claims, flagged, a person submits, amber icon. Item four, urgent care triage, a vet tech always decides, red icon.
Different clinic, different animals, same shape of answer.
The decision Rafailia would take back Every AI generated note at Yewgarden used to show up in the same plain gray box on a vet tech's screen, whether it was a reminder text or a triage flag. A busy front desk queue could not tell, at a glance, which ones were just information and which needed a fast, real decision. Badge coding each note by how much freedom it was actually given fixed that, the same lesson Driftwick learned the hard way.

Mapped onto LEAD, the shape holds. The link is what a wrong call actually costs an animal and an owner, not how polished the triage model's demo looked. The early signal is the same comparison Ilonka ran: reminders catch their own mistakes within the hour, triage mistakes might never get traced back at all. The abuse Rafailia watched for was the same one too, someone assuming a clean triage demo meant the model could take the final call. And her decision matched Ilonka's: build in the ranked order, and never hand triage the pen.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: reminders first, urgent triage never runs alone, say why in one line each.
Cost: no budget for a dedicated audit team. Whoever owns the model checks a random sample of reject or wait calls by hand every week instead of waiting on a quarterly review. Slower to build, but it closes the same gap.
The model got better, for real: say the triage model's accuracy actually catches up to a vet's own judgment. Keep it helper only anyway, because the problem was never how smart it was. It was always how long a wrong call could hide, and how little the animal could do about it.

Where people run it wrong.
They let the top ranked task's trust level leak onto the bottom one, because "the team already proved the model works."
They measure a hiring or triage model only on the easy, obvious cases, where it will always look as clean as the low stakes task next to it.
They set one confidence rule for a whole platform instead of one rule per task.

How to use it live. When an interviewer hands you four tasks to rank, buy yourself a second by asking out loud what a wrong answer actually costs on each one, in real terms, before ranking anything. That question is most of the answer already.

Flashcards (tap any card to flip it)

1 · THE FRAMEWORK
What framework fits a question that asks you to rank tasks by AI suitability?
Tap to flip
ANSWER
LEAD: link, early signal, abuse, decision. It forces "AI suitable" to mean something checkable, stakes and how fast a wrong answer gets caught, instead of a feeling.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Ilonka Vantresca, who owns the AI roadmap for Driftwick, Elderbrook's travel and expense platform, and Kalju Norrbin, Elderbrook's head of people, who asked for a fully automatic hiring model.
3 · THE LINK
What should "AI suitable" actually be measured against?
Tap to flip
ANSWER
Three things: how bad one wrong output is, how fast and cheap it is to catch, and who gets hurt if nobody catches it in time. Not the model's own accuracy score.
4 · THE RANKING
What's the actual order, most suitable to least?
Tap to flip
ANSWER
Invoice matching, expense approval, contract review, hiring decisions.
5 · THE ABUSE
How does this ranking get misused?
Tap to flip
ANSWER
Reading "high on the list" as "safe to run with zero review anywhere," instead of "safe to trust with more freedom." Kalju's version: "if invoice matching runs alone, hiring should too."
6 · THE NUMBER
Fill in the blank: Elderbrook's hiring rule ran quietly for ___ weeks before a routine audit caught it, rejecting candidates with employment gaps at ___ the rate of similar candidates.
Tap to flip
ANSWER
Eleven weeks; nearly double the rate, sixty one percent against thirty four percent. About 640 applications went through the stage, and around 60 auto rejects should have gone to a person first.
7 · THE DECISION
What does Ilonka actually do with the ranking?
Tap to flip
ANSWER
Builds in that order, invoice matching first. Keeps a person checking anything past a real dollar line even on the top ranked task. Rebuilds hiring as helper only: the model can flag, it never rejects anyone alone.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs LEAD again on a different product. Which one, and what's the parallel ranking?
Tap to flip
ANSWER
Mossglen, Yewgarden Veterinary Group's AI layer, run by Rafailia Ulfsson. There: appointment reminders rank highest, then visit summaries, then insurance claims, with urgent care triage last, where the model flags and a vet tech always decides.

Check yourself Score: 0 / 0

Multiple choice
1. Why does invoice matching rank highest among the four tasks?
  • A. It uses the newest version of the model.
  • B. A wrong match is cheap to fix and gets caught the same day.
  • C. Finance likes the tool the most out of the four teams.
  • D. It processes the most data, so it must be the most accurate.
Show hint
Check the Link and Early signal steps in the framework recap.
Show answer
B. Accuracy was never the deciding factor. A wrong invoice match costs about twelve dollars and gets caught the same business day, which is what actually earns it more freedom.
True or false
2. True or false: once the hiring model's accuracy gets as good as invoice matching's, it should be trusted to run the same way, with no review.
  • True
  • False
Show hint
Check the Abuse step. This is exactly the trap Kalju's request fell into.
Show answer
False. Accuracy was never the blocker. A wrong hiring reject can hide for a whole quarter and cannot be undone the way a twelve dollar invoice fix can, no matter how good the model's numbers look in a demo.
Fill in the blank
3. Elderbrook's hiring auto reject rule ran for ___ weeks before the quarterly audit caught it, and about ___ applications had gone through that stage by then.
Show hint
Look at the numbers in Let's learn, right after the audit is described.
Show answer
Eleven weeks; about 640 applications. Around 60 of the auto rejected candidates should have gone to a person first.
Short answer, name the reversal
4. What old decision would Ilonka take back, and why did it make sense the first time nobody thought to change it?
Show hint
Look at the block-key labeled "The choice I would take back" in Let's learn.
Show answer
Model answer: The single confidence rule, ninety percent sure, act alone, review monthly, applied to every model call on the platform. It made sense when invoice matching was the only live use case, so there was nothing to tell it apart from. It stopped making sense the moment a second, much higher stakes task got wired to the same switch.
Short answer, apply it yourself
5. Pick two tasks a tool you use today actually handles. Using how bad a wrong answer is and how fast you'd catch it, which one deserves more freedom to act alone?
Show hint
Think about cost if wrong and days to notice, not how accurate the tool seems.
Show answer
Model answer: A photo app's auto suggested album title is more AI suitable than its auto delete of "duplicate" photos. A bad title costs nothing and is obvious the moment you see it. A wrongly deleted photo might not get noticed for months, and it cannot be undone.
Short answer, work the number
6. If contract review's missed clause detection time dropped from 180 days to 14 days, say with automatic clause tracking software, would that change its place in the ranking?
Show hint
Faster verification earns more freedom, but check what happens to the stakes number too.
Show answer
It would move up, closer to expense approval. Faster verification is exactly what earns more freedom in this method. But the dollar stakes per contract, about $420,000 a year, are still far higher than expense approval's, so it likely earns a bigger flagging role, not full autonomy.
Before you close the answer
Why this works
Tests whether you rank by real stakes and how fast a wrong answer gets caught, not by a demo accuracy score or how much the person asking wants something built. It also tests whether you know a ranking sets a ceiling on freedom, not a blanket policy that applies the same way to everything on the list.
Follow-up traps
"What if the hiring model's accuracy eventually beats a human recruiter's?" Response: accuracy was never the blocker, being able to catch a wrong call fast was. A rejected candidate still has no way to contest a call in real time, and the harm still hides until a scheduled audit finds it.

"Isn't refusing to fully automate hiring just leaving efficiency on the table?" Response: the model still does the efficiency work, flagging missing skills, summarizing a resume, sorting the obvious cases. A person just keeps the one decision that cannot be undone.
If pressed
The ninety percent bar is not the same ninety percent everywhere. It is set per task, against that task's own set of already checked examples. Invoice matching's ninety percent is measured against thousands of confirmed past matches. A task with a thinner or newer set of checked examples needs a higher bar before it earns any freedom at all, not the same round number borrowed from a different job.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more