ConceptFoundationalAI Opportunity & Model Strategy / Opportunity identification for AI / #5
Explain why high-volume, low-stakes, tolerant-of-error tasks are the best first targets.
BOUND · picking Cantillion's first AI feature, an AI tool that tracks music licensing and calculates royalty payments for artists and rights holders
Cantillion pulls licensing records out of a dozen-plus databases every night and works out who gets paid what, every stream, every sync, every airplay. Wynford Bagshaw runs rights operations there. Nkiruka Athelstan runs partnerships, and she just watched Delacroix-Webb Publishing, one of Cantillion's biggest clients, spend six weeks fighting over how to split one song's royalties between three co-writers. She wants Cantillion's first AI feature to fix that. Wynford has to decide if she's right.
The direct answer
Build the AI feature where all three line up: it happens constantly, a wrong answer costs almost nothing, and the workflow already has a way to catch a miss without anyone getting hurt. That combination is what turns even a small per-item win into a real number fast, and it lets you ship before the feature is perfect. A rare, high-stakes problem might matter more per case, but you cannot learn from six cases a month, and there is often no way to even check if the model's answer is right.
Do this, in order
Build the high-volume, low-stakes, error-tolerant feature first, not the one that would impress a client.Why: this is the one decision the whole answer is built to defend.
Break down what each of the three properties actually buys you before ranking anything.Why: a task can look great on one property and fail badly on the other two.
Own the real hours with a stated assumption, not a felt sense of "this one matters more."Why: a number with its source survives a follow-up question, excitement doesn't.
Check at least four candidate tasks, not just the two obvious ones, before picking.Why: some tasks are high-volume but risky, or safe but too rare to teach you anything, and the case for going first weakens fast once only two of three properties hold.
Sanity-check the "let's lead with the exciting one" instinct against the actual arithmetic.Why: excitement optimizes for the demo. Volume and error tolerance optimize for the roadmap.
Save the rare, high-stakes task for once there's a track record and someone can build a real eval set.Why: naming when to revisit it is what makes this a sequencing call, not a rejection.
How to answer this, stage by stage
Nobody is grading whether you can recite "high volume, low stakes, error tolerant." They're grading whether you can turn those three words into a number that would actually change a real roadmap meeting.
01
Scope it to one company, one meeting, one real pitch
Say it like this
"Let's ground this. Cantillion tracks music licensing and works out royalty payments. Their head of partnerships just pitched building an AI that resolves disputed co-writer splits first, because it would win over a big publishing client. That's the specific call I'm going to make, not 'how do you prioritize AI features' in general."
Why this works
Pins an abstract prioritization question to one real, checkable choice instead of a lecture.
02
Say the rule out loud before touching a single number
Say it like this
"My test for a first AI project is three things, all at once: it happens a lot, a wrong answer costs almost nothing, and the workflow can absorb a miss without anyone getting hurt. Two out of three isn't enough yet, and I'll show you why."
Why this works
Tells the interviewer you have a real test to apply, not a preference you're about to justify after the fact.
03
Break down what each property is actually buying you, the B step
Say it like this
"High volume means even a tiny per-item win adds up fast, and it hands you a huge sample to learn from in weeks. Low stakes means a wrong output costs almost nothing, so you can ship while the model is still a bit rough. Error tolerance means the workflow can catch the miss before it does damage. Miss all three, and you're stuck trying to perfect something in the dark before you're even allowed to ship it."
Why this works
Gives the interviewer a testable reason for the rule, not just three buzzwords in a row.
04
Own the real numbers on both candidates, the O step
Say it like this
"Cantillion's catalog ingestion runs about 48,000 title-and-writer pairs a month that need matching across licensing databases, at 45 seconds of manual comparison each. Let the model propose the match with a confidence flag, and about 85 percent get a fair glance instead of a full recheck. That's roughly 419 hours saved a month. The disputed splits run about 6 cases a month, 16 hours each by hand. Automate every single one of those perfectly, which is a big if, and that's 96 hours."
Why this works
A number with a stated source survives a follow-up question. A feeling about which project matters more doesn't.
05
Show the range across more than the two obvious candidates, the U step
Say it like this
"Not every task sorts neatly. Auto-sending royalty statements straight to top artists is high volume, but it isn't low stakes, a wrong number reaching someone that big does real damage. A one-time cleanup on a small legacy catalog is low stakes, but there's barely enough volume to build a real habit or a real eval set. When only two of the three properties hold, the case for going first gets a lot weaker, not automatically dead, just weaker."
Why this works
Turns "it depends" into something checkable against real candidates instead of a shrug.
06
Run the sanity check that would have caught the wrong pitch, the N step
Say it like this
"If the exciting project saves 96 hours a month at its absolute best, and the boring one is already saving 419, picking the exciting one first isn't ambition. It's optimizing for the story you get to tell in the room instead of for the roadmap."
Why this works
This is the exact check that stops a team from shipping the wrong thing first because it played well in a meeting.
07
Name the direction, reject an alternative, and close on the real rule, the D step
Say it like this
"The single thing that would change this most isn't the size of either number, it's how often the model's confidence is actually wrong on the matching task. Even if that flag rate climbed to 80 percent, title matching would still save about as much as fully automating every dispute. So here's the rule: build the boring, high-volume, low-stakes thing first. We looked at running both at once and dropped it, splitting a small team two ways ships neither one well. Save the disputed splits for once there's a track record and a real eval set for a question that doesn't have an agreed right answer yet."
Why this works
Ends on an operating rule the interviewer can picture actually running, not a promise to be more careful.
Let's learn
What happens when a team picks its first AI project by which one gets the loudest applause in a room, instead of by the numbers?
Say a company builds a tool that reads licensing records from a dozen databases and works out who gets paid, every time a song streams, airs on the radio, or gets used in an ad. One of the jobs that tool has to do constantly, quietly, in the background, is match a new catalog record against records it already knows, because two databases almost never spell a song title or a writer credit exactly the same way.
The whole feature in five boxes. The fourth one, the flagged full check, is the reason a wrong guess never turns into a wrong payment.
Before any AI touched this, an analyst compared each pair by hand, title against title, writer against writer, duration against duration, about 45 seconds a pair. With roughly 48,000 pairs landing every month, that was 600 hours of pure comparison work a month, spread across a small team.
With the model proposing each match and flagging how sure it is, the sure ones get a ten-second glance instead of a full compare, and the unsure ones still get the same full 45-second check as always, nothing skipped. That drops the monthly total to about 181 hours. A real save of 419 hours a month.
Knowledge spark: what's a confidence flag?
A number the model attaches to its own guess, how sure it is that this match is right. High means a quick glance is safe. Low means a person needs to actually look. It turns "trust the whole thing or check the whole thing" into a dial the team can set on purpose.
Here's the turn. The thing that got everyone excited in a roadmap meeting wasn't this. It was something messier: teaching the model to work out which of three songwriters gets which cut of a disputed song's royalties, all on its own. Say plainly what's true: that excitement is not the problem. Chasing it first, before there's any track record, is.
Ninety-six hours was the ceiling on the exciting project. Four hundred and nineteen was the floor on the boring one.
What it costs at its worst: the team spends a full quarter building an autonomous split-resolver nobody can actually check, because the "right" split is the exact thing being disputed, there is no ground truth to test it against. If it ships anyway and gets one call wrong, in a way that pays real money to the wrong person, that's worse than never building either feature. Now a client is disputing the AI's answer on top of disputing the original songwriting credit. Meanwhile the 48,000-pair-a-month backlog keeps eating hours nobody touched.
The choice I would take back
Cantillion's roadmap review let whichever client complaint felt biggest and most visible set what got built next. That was fine back when the backlog was small and every candidate project was roughly the same size. It stopped being fine the moment "biggest and most visible" quietly stopped matching "biggest actual return."
What I would leave alone: this isn't an argument for never building the disputed-split feature. Some rare, high-stakes problems really are worth solving first, when there's no safe workaround and someone is walking away today. This wasn't that. Cantillion could keep resolving disputes by hand for another two quarters without anyone getting hurt, so there was no honest reason to rush it.
The lesson: don't let the size of the reaction in the room stand in for the size of the number. The task nobody claps for in a pitch meeting can still be the one worth 400 hours a month.
Now here is the same thing as a story
The short version is above, for saying out loud. Read this one for the actual Thursday the two numbers finally sat in the same room.
Wynford Bagshaw can tell, from three lines of a licensing feed, whether two records describe the same song or two different ones that just happen to look alike. He's run rights operations at Cantillion for four years, and the catalog-matching backlog has been his job, unglamorous and constant, the whole time.
Nkiruka Athelstan is good at a different thing: a room. She runs partnerships, and she'd just come off a call with Delacroix-Webb Publishing, a mid-size independent publisher and one of Cantillion's oldest clients. Three songwriters on one Delacroix-Webb track had spent six weeks arguing over the royalty split, each one certain of a different number, no contract clean enough to settle it outright. Delacroix-Webb's founder had said, almost as a joke, "can't your AI just tell us who's right?"
Nkiruka did not think it was a joke. At the next roadmap review, she pitched it as Cantillion's first real AI investment: an autonomous dispute-resolution feature, something to show every publisher client. The room liked it. Rooms like a good story.
The two candidates as they actually looked, side by side, before anyone had run a single number.
Wynford was asked to scope it. He started, as he always did, by trying to write down what "right" would even mean. For the matching backlog, right meant matching Cantillion's own master catalog, a real answer already sitting in the database. For the disputed split, there was no master answer. The correct percentage was the exact thing three people were shouting about. He couldn't build an eval set for a question nobody agreed had a right answer.
Knowledge spark: what's an eval set?
A pile of real examples with a known right answer, used to check how often a model actually gets it right before anyone trusts it. No known right answer means no eval set, and no eval set means nobody can honestly say the model is ready.
So instead he ran the numbers he could actually check. Forty-eight thousand pairs a month, 45 seconds each by hand, 600 hours. With a confidence flag doing the sorting, 85 percent get an eight-second glance, 15 percent still get the full 45-second check. That's about 181 hours. A save of 419 hours, every single month, starting the month it ships.
Then the disputed splits: six cases a month, sixteen hours each of contracts, session logs, and old statements. Automate every one of them, perfectly, no mistakes, no walk-backs. That ceiling is 96 hours.
Four real candidates, plotted honestly. Only one of them sits in the corner where a first AI project actually wants to be.
He brought both numbers to the next review, on one slide, no framing, just the arithmetic. Nkiruka looked at it longer than anyone else in the room. She hadn't been wrong that the dispute mattered, Delacroix-Webb's frustration was real. She'd just been ranking by how loud a problem sounded, not by how much it actually moved.
The same four candidates, this time sorted by the actual decision each one earns.
"We're not saying no to the dispute feature," Wynford told her. "We're saying not yet, and not alone. It needs a way to check itself, and right now it doesn't have one."
Six weeks, start to first ship. The loud project didn't disappear. It just stopped going first.
Title matching shipped in week eight. By the end of its first full quarter, it had saved rights operations just over 1,250 hours, more than one analyst's entire year, and every royalty it touched had still passed through a human glance or a human check, nothing paid out unattended. The disputed-split feature is still on the roadmap, redesigned now as a tool that hands a specialist the contracts, the session logs, and a first-pass draft split, and lets a person make the actual call.
What Wynford would tell his past self: he'd let "how big does this feel in the room" quietly stand in for "how big is this, actually," for longer than he'd like to admit. The task nobody was excited about in that first meeting was the one that was already worth more than the one everyone wanted to talk about.
BOUND, for choosing which royalty problem earns Cantillion's first model
Not a way to make an exciting idea sound suspicious. BOUND turns "this feels like the right first project" into a number Wynford, or anyone else at Cantillion, could actually defend in the room.
BBreak it down. What does each property actually buy you?
High volume means a small per-item win multiplies into a real number, and it hands you a big real sample to learn from in weeks, not a quarter. Low stakes means a wrong output costs almost nothing, so the team can ship before the model is perfect instead of after. Error tolerance means the workflow, or a person in it, can absorb a miss without real harm. All three have to hold, not just one that sounds good in a pitch.
Skip this split and "good first AI project" stays a feeling instead of a checkable list of three separate things.
Three properties, one gauge. Miss any one of them and the needle doesn't actually point at a safe first project.
OOwn the numbers. Where does each one actually come from?
Title matching: 48,000 pairs a month, 45 seconds manual, 600 hours old way. With an 85/15 confident/flagged split, 8 seconds for a glance and 45 seconds for a full check, the new total is about 181 hours. Saved: 419 hours a month. Disputed splits: 6 cases a month, 16 hours each by hand, so the absolute ceiling if AI resolved every single one perfectly is 96 hours a month.
A number only counts as owned if you can say exactly where it came from when someone pushes on it in the room.
Where 419 hours a month of saved time actually comes from
Old, fully manualNew: flagged pairs, still a full checkNew: confident pairs, quick glance
Every hour saved comes from the confident 85 percent. The flagged 15 percent costs exactly what it always did, on purpose.
UUse a range, not one flattering number.
Four real candidates, not two: title matching, high volume, low stakes, the corner to start in. Disputed splits, low volume, high stakes, no eval set possible. Auto-sending royalty statements straight to top artists, high volume but high stakes, a wrong number reaching someone that size does real damage. A one-time cleanup on a small legacy catalog, low stakes but too low volume to teach the team much of anything.
A single "build the AI thing" answer here repeats Nkiruka's exact mistake, treating four very different bets as if they were interchangeable.
Four separate reasons this specific task can absorb a wrong guess without anyone getting hurt.
NNail the sanity check. Does the exciting pitch survive contact with the real numbers?
If the project everyone got excited about saves 96 hours a month at its absolute best, and the boring one is already saving 419, building the exciting one first isn't ambition. It's optimizing for the reaction in the room instead of the actual roadmap.
This is the exact check that would have stopped the pitch before it ate a quarter nobody could get back.
How far the flagged-review rate would have to climb before this stops winning
Hours saved, title matchingDisputed-split ceiling
The model's flag rate would have to more than quintuple, from 15 percent to about 80 percent, before this stops beating even a perfect version of the other project.
DDirection. Which assumption moves this most, and what's the actual decision?
Not the size of either headline number. The single biggest swing factor is how often the model's own confidence is wrong on the matching task, and the chart above shows that factor has enormous room before it changes the answer. So the real decision is a sequencing rule, not a rejection: build the high-volume, low-stakes, error-tolerant feature first, always, and revisit a rare, high-stakes feature once there's a track record and a real way to check it.
Naming the fact that actually swings the estimate, instead of the biggest number in the room, is what separates a real estimator from a confident guesser.
One alternative considered and rejected: running both projects in parallel, on split teams, so nobody had to say no to Nkiruka's pitch. It lost, because a small rights-ops team split two ways ships neither project well, and title matching's 419 hours a month were sitting there, unclaimed, for every week that decision sat unmade. The AI-specific risk underneath all of this is quiet and easy to miss: a model can confidently propose a match between two records that only look alike, in the exact same tone as a correct match, with nothing on the screen to flag the difference. The guardrail is the confidence threshold itself, calibrated against real held-out pairs, plus the mandatory full check on anything it flags, so a wrong guess never reaches a payment unattended. And the trade-off is real, not free: even the confident 85 percent still gets a human glance, not a fully hands-off pass, trading a little speed and cost for a workflow that can catch its own mistakes.
And if you want to be sure it really works, try it somewhere else
Same five letters, a library system instead of a rights ledger, and this time the tempting rare-and-risky project is an insurance appraisal instead of a royalty split.
Spinecheck matches book records across the branches of the Fenlake Library Consortium, so a hold or a fine follows the right physical copy instead of stalling on a title that's spelled two different ways in two different systems. Undine Thorsby leads cataloguing there, and she ran into the same choice Wynford did, on a much smaller shelf.
Run BOUND on it. Break it down: matching near-duplicate book records happens constantly, across a busy consortium, and a wrong link just delays one hold by a day, easily caught and fixed. Appraising a rare book's insurance value after fire or flood damage happens rarely, maybe four times a year, and a wrong number changes what an insurer actually pays out on a claim, hard to reverse once it's filed. Own the numbers: Fenlake's branches generate about 14,000 candidate record-matches a month, at roughly 30 seconds of manual comparison each, 116.7 hours old way. A confidence-flagged match, glance-confirmed on the easy 90 percent, drops that to about 41 hours, a save of 75.7 hours a month. Four disputed appraisals a year, at about 20 hours of expert review each, is a ceiling of 80 hours a year, or well under 7 hours a month.
A different shelf, the exact same shape of choice: a lot of cheap misses, or a few expensive ones with no real way to check them yet.
Where Fenlake's answer genuinely differs
Undine's version has an even bigger gap, 75.7 hours a month against under 7, because rare-book appraisal disputes are rarer than co-writer splits ever were. But the reasoning underneath is identical: volume compounds, a single high-stakes case doesn't, and there's no honest eval set for a value that's disputed precisely because nobody agrees on it yet.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Name the three properties, own one real number, and say plainly that a range beats a single flattering guess.
Cost: no time to run a full comparison before the pitch. Shrink the sample to the handful of cases that would actually decide the argument, and say honestly that it's a smaller sample.
The model got better, for real: a new release makes the confidence flag far more reliable. Recheck the flag rate anyway, because a better model can still shift which task deserves to go first next, not just how fast the current one runs.
Where people run it wrong.
They rank by how much a single case matters instead of by how often it happens.
They assume "we can automate it eventually" means "we should try to automate it first."
They treat a rare, high-stakes task as unsolvable forever, instead of just wrong to build first, before there's a track record.
How to use it live. Before answering, ask yourself one plain question out loud: "does this happen constantly, does a miss cost almost nothing, and can the workflow catch it?" If you can only say yes to two of the three, say which two, out loud, and let that be the honest hedge instead of a guess dressed up as confidence.
Flashcards (tap any card to flip it)
1 · THE METHOD
What method fits picking which of two AI features a team should build first?
Tap to flip
ANSWER
BOUND: break it down, own the numbers, use a range, nail the sanity check, name the direction. Built for turning a gut feeling about priority into arithmetic you can defend.
2 · WHO'S IN IT
Who is this answer about?
Tap to flip
ANSWER
Wynford Bagshaw, Cantillion's rights-ops lead. Nkiruka Athelstan, Cantillion's head of partnerships. Delacroix-Webb Publishing, the client whose disputed co-writer split started the whole pitch.
3 · THE THREE PROPERTIES
What three things does a good first AI project need, all at once?
Tap to flip
ANSWER
High volume, so a small win compounds fast. Low stakes, so a wrong output costs almost nothing. Error tolerance, so the workflow can absorb a miss without real harm.
4 · THE REAL HOURS
Fill in the blank: title matching saves about ___ hours a month. Fully automating every disputed split, perfectly, would save at most about ___ hours a month.
Tap to flip
ANSWER
419 hours a month. 96 hours a month. Same company, same quarter, more than four times the return from the boring project.
5 · WHEN TWO OF THREE HOLD
Name a task that's high volume but NOT low stakes. What happens to its case for going first?
Tap to flip
ANSWER
Auto-sending royalty statements straight to top artists with no review. It happens constantly, but a wrong number reaching someone that size does real damage, so the case for going first gets much weaker, even though volume alone looks great.
6 · THE GUT CHECK
Why was Nkiruka's pitch to build the disputed-split feature first the wrong call, even though the underlying problem was real?
Tap to flip
ANSWER
She ranked by how loud the problem sounded, not by how much it actually moved. Ninety-six hours a month, at best, against 419 already sitting unclaimed in the matching backlog.
7 · WHAT WOULD CHANGE THIS
Which single assumption swings this estimate most, and how much room does it have before the answer changes?
Tap to flip
ANSWER
How often the model's own confidence is wrong, the flagged-review rate. It would have to climb from 15 percent to about 80 percent before title matching stopped beating even a flawless disputed-split feature.
8 · SAME METHOD ELSEWHERE
Section 4 runs BOUND again on a different product. Which one, and what's the same shape?
Tap to flip
ANSWER
Spinecheck, a library-catalogue matching tool for the Fenlake Library Consortium. Same shape: constant, cheap book-record matches beat rare, high-stakes rare-book appraisal disputes by an even wider margin, 75.7 hours a month against under 7.
Check yourself Score: 0 / 0
True or false
1. True or false: fully automating every disputed co-writer split, perfectly, with zero mistakes, would still save more time each month than the AI-assisted title matching project.
True
False
Show hint
Check the O step's two owned numbers.
Show answer
False. Even the best-case ceiling for disputed splits, 96 hours a month, is less than a quarter of what title matching already saves, 419 hours, with real, non-perfect assumptions.
Fill in the blank
2. Cantillion's title-matching backlog runs about ___ pairs a month. Matching it with an AI-assisted confidence flag instead of a full manual check saves about ___ hours a month.
Show hint
Check the O step and the stacked bar chart.
Show answer
48,000 pairs. 419 hours. The saving comes entirely from the 85 percent the model is confident about, glanced at instead of fully rechecked.
Multiple choice
3. Why does the disputed co-writer split project make a worse first AI project, even though a single case matters more in real money?
A. Cantillion's engineers don't have the skills to build it yet.
B. It happens too rarely to teach the team much, and there's no ground truth to check the model's answer against.
C. Delacroix-Webb Publishing refused to let Cantillion work on their case.
D. The model can't process contracts and session logs at all.
Show hint
Check the B and U steps.
Show answer
B. Low volume means it can't compound into a real number fast, and the "right" split being disputed means there's no honest eval set to check the model against.
Short answer, where it would be worth it anyway
4. Name a case where Cantillion genuinely should build a rare, high-stakes AI feature first, skipping the volume test entirely.
Show hint
Check "what I would leave alone" in Let's learn.
Show answer
Model answer: If a client were about to walk away today over exactly this problem, with no safe workaround available, the urgency itself would outweigh the usual volume test. That wasn't Delacroix-Webb's situation, they could wait two more quarters without real harm.
Short answer, apply it yourself
5. Think of a task you or your team does often that's tedious and low-stakes. What's the real monthly time it would save if AI took even a modest, imperfect bite out of it?
Show hint
Estimate a volume, a per-item time, and a realistic (not perfect) share the model could handle safely.
Show answer
Model answer: Sorting a shared inbox by topic before routing it. Say 2,000 emails a month, 20 seconds each to read and route by hand, 11 hours a month. Even a modest model that gets 70 percent of them right, glance-confirmed, could cut that closer to 4 hours, a small number that still beats chasing one flashy, rare automation elsewhere.
Short answer, work the number
6. If the model's flagged-review rate on title matching rose from 15 percent to 40 percent, would it still save more time a month than fully automating every disputed split? Show the arithmetic.
Show hint
Use the line chart's formula: hours saved equals 493.3 times (1 minus the flagged rate).
Show answer
Yes. 493.3 times (1 minus 0.40) equals about 296 hours a month, still well above the disputed-split ceiling of 96 hours, even with the flag rate more than doubled from where it actually sits today.
Before you close the answer
Why this works
Tests whether a candidate can resist the pull of a rare, executive-visible problem and instead defend a boring choice with real arithmetic, when it's a team's first AI investment and the wrong call is expensive to walk back.
Follow-up traps
"Isn't one big dispute worth more, case for case, than one small routing delay?" Response: per case, sure, but Cantillion isn't paid case by case, it's paid in hours freed up company-wide, and 48,000 small wins beat 6 big ones because the aggregate compounds and the big ones don't.
"What if there's a real regulatory deadline forcing the dispute feature to ship first?" Response: then it's not really a first-choice decision anymore, it's a forced dependency, and the honest move is to say so plainly rather than pretend the volume math still picked it.
If pressed
The confidence threshold that sorts the 85/15 split sits at a model score of about 0.90. That number wasn't a guess, pairs scoring between 0.80 and 0.90 in the calibration set turned out to have a real mismatch about one time in six, which is exactly the group that still gets the full 45-second check.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.