Rank these by AI suitability and justify: expense approval, contract review, invoice matching, hiring decisions.
Driftwick is Elderbrook's travel and expense platform. It books flights and hotels, matches travel invoices to the trip they belong to, and drafts expense reports so nobody has to open a spreadsheet by hand. Ilonka Vantresca owns the small AI team that builds all of that. This quarter, four different requests landed on her roadmap at once: keep matching invoices, start approving expense reports, start flagging risky lines in hotel contracts, and, from Kalju Norrbin, Elderbrook's head of people, help decide who gets hired. Only one of those four was ever going to be safe to hand over completely.
- Rank invoice matching, expense approval, contract review, then hiring decisions, in that order.Why: this is the direct answer, just laid out as a list.
- Let invoice matching run itself under a real dollar line.Why: a wrong one is cheap to fix and gets caught the same day, so it earns the most freedom.
- Keep expense approval assisted: the model approves, a manager still sees anything unusual.Why: real money and trust sit on this one, and policy rules alone miss the odd real case.
- Let contract review flag. Never let it sign.Why: one missed clause can cost hundreds of thousands, and it might not get noticed for six months.
- Keep hiring fully human. The model can point out a missing skill. It never decides.Why: the person harmed by a wrong call can't push back, and the mistake can hide for a whole quarter.
- Kill the one confidence rule running every model call the same way. Give each task its own.Why: ninety percent sure means something different on a twelve dollar invoice than it does on somebody's job.
How to answer this, stage by stage
Nobody is grading whether you can name four tasks in order. They're grading whether you can say, out loud, what actually earns a task a spot on this list.
Let's learn
What happens when four different teams ask the same small AI group to automate four completely different jobs, all at once?
Driftwick is Elderbrook's travel and expense platform. Book a flight, book a hotel, and Driftwick's AI matches the invoice to the trip it belongs to and drafts the expense report before anyone has to open a spreadsheet.
When Driftwick's AI first went live, invoice matching was the only thing running. Ilonka's team wrote one rule for it: if the model was more than ninety percent sure, and the amount was small enough that a mistake wouldn't matter, let it act alone. Check everything else once a month. That rule automatically matched 38,000 invoices a month, and it saved the finance team about three hundred hours of manual matching every month, real hours, given back.
Over the next year, three more requests landed on that same rule without anyone sitting down to ask whether it still fit. Expense approval got wired to it in spring. A first pass at flagging risky contract clauses got wired to it in summer. And, quietly, in September, Kalju Norrbin's hiring pipeline got an auto reject step built on the exact same rule too, to clear a three week backlog of applications before a hiring freeze made the delay somebody else's problem.
None of the four broke because the model got worse. That's not the story here. Ninety percent sure, act alone, means something different depending on what sits behind it. On an invoice it means fixing a twelve dollar typo alone. On a hiring decision it means ending someone's application alone, and not finding out for months if it was even the right call.
Line up the real numbers behind those four dots. Invoice matching: about 38,000 a month, a wrong one costs about twelve dollars to fix, and finance sees it the same business day, because the match either lines up against the booking record or it doesn't. Expense approval: a wrong auto approval costs a few hundred dollars, and a weekly manager review catches it inside about four days. Contract review: Elderbrook signs or renews about sixty hotel and vendor rate deals a year, each worth around $420,000 in yearly spend, and a missed clause can sit unnoticed for about six months, roughly 180 days, until a hotel invokes it or a renewal date quietly passes. Last year, a missed auto renewal clause on one hotel block locked Elderbrook into an old rate for another twelve months once the window to renegotiate had already closed. Real cost: about $180,000.
Hiring sits furthest from safe, and here is why. The check that would catch a bad call, Elderbrook's routine equal opportunity audit, only runs once a quarter, about ninety days between checks. This time it happened to land at eleven weeks, seventy seven days, a bit of luck in the calendar, not something the design earned. The audit found that candidates with a gap in their work history were being rejected at nearly double the rate of otherwise similar candidates, sixty one percent against thirty four percent. Employment gaps track with caregiving leave. Caregiving leave tracks, more than most places want to admit, with gender.
About 640 applications went through that stage in eleven weeks. Around sixty of the auto rejected ones should have gone to a person first. Untangling it, the legal review, reopening the cases, rebuilding the process, ran to about $340,000, and that is before counting what it costs a company to sit under an equal opportunity finding.
What I would leave alone: Driftwick's AI also decides which hotels and flights show up first in a traveler's search results, with zero review, and that is exactly right. A ranked list is a suggestion. Nobody is harmed by seeing the third best hotel before the best one, they just scroll past it. Being wrong there costs nothing, because nothing was ever decided, only offered.
The lesson: "AI suitable" is not one number set once for a whole platform. It is a separate call for every task, about how bad a wrong answer is and how fast anyone would know. Treat every model call as the same call, and a genuinely useful invoice tool ends up quietly deciding something about somebody's job that nobody meant it to decide.
Now here is the same thing as a story
Use this version when you have the time. The short version above is what you would actually say out loud in the room. This one is for feeling why the order matters.
Ilonka can tell which AI request is actually simple, and which one is a lawsuit with a nice screen on top, usually before the person asking finishes their sentence. She has run Driftwick's AI team for three years, and she built the very first version of the platform's one automated rule herself: read a hotel or airline invoice, match it against the trip it belongs to, let the match through alone if the model is more than ninety percent sure and the amount is small enough that a mistake would not matter. It was the only automated thing on the whole platform, and it worked so well that finance stopped opening the matching queue most mornings.
For a year, that felt like the whole job. New requests kept landing, and each one got quietly pointed at the same rule, because the rule already existed and it already worked. Nobody sat down and asked whether ninety percent sure, act alone, should mean the same thing for a fifty dollar taxi receipt as it does for a four hundred thousand dollar hotel contract. It was just the rule. It was right there.
Then, one afternoon in September, Kalju Norrbin came by with a backlog problem. Elderbrook's own early stage resume screen was three weeks behind, and a hiring freeze was about to make that delay somebody else's fault instead of his. Ilonka's team had the rule sitting right there, already tested, already running three other jobs. It took an afternoon to point it at resumes instead of receipts. Nobody meant any harm by it. It was the fastest fix in the building.
For eleven weeks it ran quietly, doing exactly what it was told. Nobody watched it the way finance watched the invoice queue, because on the screen, nothing about it looked different from the other three jobs already running the same rule.
Then the quarterly audit came around, the way it always does, checking a random sample of hiring decisions for basic fairness. This time the sample flagged something real: candidates with a gap in their work history were being auto rejected at nearly double the rate of otherwise similar candidates. Nobody had done anything wrong on purpose. The model had learned, from years of past hiring data, that gaps tracked weakly with worse outcomes down the line, and ninety percent sure turned that weak pattern into reject, no review, every time.
Kalju, to his credit, did not try to talk his way around the number once it was in front of him. His question was the honest one: "If the model's good enough to run invoice matching alone, why can't it run this alone too, once we fix the gap issue?" Ilonka told him no, and had to say why more than once. Being good enough to run alone on an invoice was never really about how smart the model was. It was about a wrong invoice costing twelve dollars and getting caught the same afternoon. A wrong reject cost somebody a chance they never knew they'd lost, and the only thing checking for it ran four times a year.
Run the same eleven weeks again, with the new rule instead of the old one. The model still flags the same six hundred forty resumes, for the same reasons. But every reject now goes to a person before anyone's application actually closes. The worst case stops being sixty people rejected wrongly and found three months late. It becomes zero, because the model was never holding the pen. It was only pointing at the page.
What I'd tell myself, back when that first rule felt so clean it seemed almost wasteful not to reuse it: the rule was never wrong. It was right, for invoices. Reusing it everywhere was not being efficient. It was assuming four different jobs were the same job, because they all happened to run through the same kind of model.
LEAD, so "AI suitable" means something you can check, not a feeling
Not a way to prove a model is smart enough. LEAD forces you to say exactly what a wrong answer costs, and exactly how fast you would know, before anything gets to run alone.
The recap, one line per letter: link "AI suitable" to how bad a wrong answer is, how fast anyone would catch it, and who gets hurt, not to a demo accuracy score. The early signal is the actual comparison across the four tasks, and it is why invoice matching leads and hiring trails, by a wide margin. Name the abuse plainly: treating a high rank as permission to skip review everywhere, and building an eval set that only ever shows the easy cases. And the decision is what makes any of this real: build in the ranked order, and keep a person in the loop at the top of the list too, not just the bottom.
Two things worth naming outright, since the real judgment sits here. Kalju's ask, to let the hiring model decide once it "got good enough," was taken seriously and turned down, not because the team distrusted the model's numbers, but because a rejected candidate has no way to ever push back on a call they never even knew got made, and a scheduled quarterly audit is the fastest anyone would notice something went wrong. The AI specific failure worth naming by name is exactly what the audit caught: a fairness gap that never shows up in an aggregate accuracy number, only in a check broken out by group, run on a schedule, not waited on. And the trade off was real and named on purpose: Ilonka could raise invoice matching's bar from ninety to ninety eight percent sure, making the auto run rule even safer, but that pushes thousands more invoices into the manual queue every month, trading a rare twelve dollar mistake for a very real chunk of finance's week. Ninety percent, under fifty dollars, stayed the line, because on this specific task, that is where the math actually holds.
And if you want to be sure it really works, try it somewhere else
Same four letters, a waiting room instead of a departures board, and this time the highest stakes task is about an animal that cannot say what is wrong.
Yewgarden Veterinary Group runs eleven clinics, and its AI layer, Mossglen, handles four requests of its own. Rafailia Ulfsson, the group's operations director, got asked to rank the same way Ilonka did: appointment reminders sent automatically, visit note summaries turned from a vet's spoken record into a structured chart, insurance claims auto filled and submitted, and phone in triage, deciding which symptoms need to be seen today and which can wait for a scheduled visit.
Reminders rank first: about 12,000 sent a month, next to no cost if one is wrong, since a client just calls to reschedule, and the mistake is obvious within the hour. Visit summaries come next: about 3,000 visits a month, a wrong summary might mean a missed detail costing around $150 in a repeat exam, caught within about a week when the vet reviews the chart before a follow up. Insurance claims sit third: about 800 a month, a wrong submission risks a denial or an overpayment worth about $220, caught in three to four weeks once the insurer processes it. Triage sits last, by a wide margin: the stakes are an animal's health, the model cannot ask a follow up question the way a vet tech would, and a wrong "it can wait" call might not get noticed for weeks, if it ever gets traced back to the call at all. The pet owner it fails has no way to contest a decision made over the phone in real time.
Mapped onto LEAD, the shape holds. The link is what a wrong call actually costs an animal and an owner, not how polished the triage model's demo looked. The early signal is the same comparison Ilonka ran: reminders catch their own mistakes within the hour, triage mistakes might never get traced back at all. The abuse Rafailia watched for was the same one too, someone assuming a clean triage demo meant the model could take the final call. And her decision matched Ilonka's: build in the ranked order, and never hand triage the pen.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: reminders first, urgent triage never runs alone, say why in one line each.
Cost: no budget for a dedicated audit team. Whoever owns the model checks a random sample of reject or wait calls by hand every week instead of waiting on a quarterly review. Slower to build, but it closes the same gap.
The model got better, for real: say the triage model's accuracy actually catches up to a vet's own judgment. Keep it helper only anyway, because the problem was never how smart it was. It was always how long a wrong call could hide, and how little the animal could do about it.
Where people run it wrong.
They let the top ranked task's trust level leak onto the bottom one, because "the team already proved the model works."
They measure a hiring or triage model only on the easy, obvious cases, where it will always look as clean as the low stakes task next to it.
They set one confidence rule for a whole platform instead of one rule per task.
How to use it live. When an interviewer hands you four tasks to rank, buy yourself a second by asking out loud what a wrong answer actually costs on each one, in real terms, before ranking anything. That question is most of the answer already.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"Isn't refusing to fully automate hiring just leaving efficiency on the table?" Response: the model still does the efficiency work, flagging missing skills, summarizing a resume, sorting the obvious cases. A person just keeps the one decision that cannot be undone.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Opportunity identification for AI
- #1 What characteristics make a workflow a good candidate for AI? List five.
- #2 Describe a method for finding AI opportunities inside an existing product without starting from the technology.
- #3 How do you distinguish a problem AI solves from a problem AI merely touches?
- #5 Explain why high-volume, low-stakes, tolerant-of-error tasks are the best first targets.
- #6 Your support team handles 8,000 tickets a month. Structure a discovery process to find the AI opportunity.
- #7 What signals in user research suggest an AI solution rather than a better interface?