ConceptFoundationalAI Opportunity & Model Strategy / Opportunity identification for AI / #1
What characteristics make a workflow a good candidate for AI? List five.
ORDER · what earns a workflow a place inside Drayline's model
Sarsen Freight Co. runs regional trucking and last-mile delivery, and its dispatch team leans on Drayline, an AI tool that reads a day's stops and the live traffic and hands each driver the order to run them in. It works well enough that new client contracts get sold with a promise: Drayline routes it from day one. Product lead Orlagh Enright used to say yes to whichever new lane leadership got most excited about that week. One of those lanes was a hospital pharmacy contract, and it very nearly went wrong before anyone checked whether the lane was even ready for a model at all.
The direct answer
Rank a workflow as an AI candidate only after it clears five checks, in this order of weight: enough real past examples to learn from, a fast and cheap way to tell if an output was right, a clear line for what counts as good enough, room to be wrong sometimes without real harm, and enough repeats to be worth the build. Fail the first check and it is not a candidate yet, however good it looks in a sales pitch.
Do this, in order
Rank a workflow only after it clears the data check, the feedback check, the good-enough check, the harm check, and the volume check, in that order.Why: this is the actual order of weight, not a wish list of nice-to-haves.
Check for enough real, representative past examples before anything else.Why: with no real history, a model has nothing to learn the pattern from, whatever the team building it can do.
Build in a fast, cheap way to tell if an output was right.Why: slow feedback means a mistake repeats for weeks before the pattern shows up anywhere.
Write down what "good enough" means before the workflow gets built, not after.Why: with no stated bar, nobody can tell a working version from a broken one.
Make sure a wrong answer here is cheap, not severe, before handing it full autonomy.Why: a high-stakes workflow needs a person in the loop until the model has actually earned trust.
Weigh volume last, and never let it excuse skipping the first four.Why: a workflow that happens a lot but has no real data behind it is still not ready, whatever the payoff looks like on paper.
How to answer this, stage by stage
Nobody is grading whether Orlagh can name five nice traits for five minutes. They're grading whether the order she puts them in would actually survive a lane that looks perfect and isn't.
1
Pin the question to one product, one lane, one real decision
Say it like this
"Let's make this real. I run product for Drayline, the route planner at Sarsen Freight Co. Say a new client signs: a hospital pharmacy network wants their refrigerated deliveries on it, and sales already told them it starts day one."
Why this works
Naming a real product and a real lane keeps the answer from turning into a lecture on prioritization in general.
2
Say your structure out loud before diving in
Say it like this
"I'd run this as ORDER. Outcome, what I'm protecting. Reversibility, which mistake is hardest to undo. Dependency, what has to be true before a workflow even counts as a candidate. Evidence, what's cheap to check first. Rank, the actual five things I'm weighing, stated in order."
Why this works
Two seconds of structure tells the interviewer a method is coming, not just an opinion.
3
Reframe what the question is really testing
Say it like this
"This isn't really 'what makes a nice AI feature.' It's asking whether I can tell a workflow that's ready for a model apart from one that only sounds ready, before I spend a sprint building it."
Why this works
Separates a real answer from a generic list of "good project" traits that could apply to anything.
4
Give the ranked five, plainly, before any evidence
Say it like this
"Here's what I'd actually check, in order. Enough real past examples to learn from. A fast, cheap way to tell if an output was right. A clear line for good enough. Room to be wrong sometimes without real harm. And enough repeats to be worth the build. Fail the first one and nothing else on this list matters yet."
Why this works
This is the direct answer, said out loud, before a single story arrives to back it up.
5
Prove it with the failure, compressed to four sentences
Say it like this
"Here's what happens without this. We put a hospital cold-chain contract on full autonomy from day one, because sales had already promised it. The lane had six weeks of local history behind it, 260 deliveries. A closed lane the model had never seen ate 53 minutes out of a 240-minute cold-chain window, and we came within 22 minutes of losing a shipment of vaccines."
Why this works
Shows the real cost of skipping the data check in one breath, not a five-minute story.
6
Name the option you turned down
Say it like this
"We could have paused all new-lane automation for a full quarter, company-wide. I'd turn that down. It throws away a routing engine that's genuinely good on the lanes it already knows. The honest fix is advisory mode: Drayline still suggests the route, a person confirms it, until the lane earns full trust."
Why this works
Shows a real judgment call happened, not just the path that shipped by default.
7
Say what you'd leave alone
Say it like this
"The established dry-goods lanes don't need any of this. Over 300,000 deliveries behind them, four years of history, a 91% match rate against what drivers actually found best themselves. Full autonomy there is earned, not assumed."
Why this works
Shows judgment instead of blanket caution applied everywhere out of fear.
8
Close on the rank, in one line
Say it like this
"Data first, feedback speed second, a stated good-enough bar third, tolerance for being wrong fourth, volume last. Any new lane sits in advisory mode until it clears the floor. That's the whole system."
Why this works
Ends on something the interviewer can actually hold the candidate to later, not just a confident sign-off.
Let's learn
Drayline is the model that looks at a day's stops, the trucks on hand, and the live traffic, then hands each driver the order to run their stops in.
Before Drayline, a Sarsen dispatcher planned a 48-truck shift by hand each morning: matching stops to trucks, checking traffic, working out an order, about 85 minutes of pencil work before the first truck could even leave the yard. Drayline cut that to roughly 10 minutes of review, and it earned that trust the slow way: four years live, over 300,000 completed deliveries behind its core model, a 91% match rate against what a driver's own best route would have found on their own.
What does "match rate" mean here?
How often Drayline's suggested route lines up with the route a driver actually found best, checked after the fact against real delivery times. A high match rate means the model has genuinely learned the lane. A low one means it's mostly guessing.
So when a hospital pharmacy network signed on for refrigerated last-mile delivery, sales promised them the same thing every other client got: Drayline routes it from day one. Nobody asked whether a six-week-old lane had anything close to 300,000 deliveries behind it. It had 260.
Route match rate: established lanes vs. the six-week-old lane
A 33-point gap in match rate is not a rounding error. It is what a model looks like when it has almost nothing real to learn a lane from yet.
The check that never ran before the pharmacy lane went live on full autonomy.
Here is the turn. A lower match rate on a brand-new lane was never really the problem. The real problem was what the model did the first time it hit a road it had never actually seen.
The pharmacy lane's cold-chain window runs 240 minutes from pickup to hospital dock, the full stretch of time the insulated boxes can hold their temperature. A normal run on that lane takes about 165 minutes, leaving 75 minutes of buffer to spare. On a Tuesday six weeks in, Drayline sent driver Ewart down a lane that had been closed for construction for ten days, a closure it had never once seen in its own thin local history. The detour added 53 minutes. Transit time hit 218 minutes. Buffer at the dock: 22 minutes.
The cold-chain buffer, draining minute by minute
Buffer, normal paceBuffer, after the wrong turn
The rust segment is the 53 minutes the closed lane cost. It's the difference between a routine delivery and a near miss.
We didn't lose 33 points of match rate. We nearly lost a shipment of vaccines that couldn't wait.
The choice I would take back
We made full autonomy the default setting for any newly signed lane, the moment it was signed, because that was the pitch sales used to win the contract. I would flip that default: a new lane starts in advisory mode, and only earns full autonomy once it clears a real data floor.
At its worst, a workflow with too little local history doesn't fail loudly. Drayline showed Ewart one calm blue line on the dash, with the exact same confidence it shows on a lane it has run 300,000 times. Nothing on the screen told him this one was closer to a guess.
What I would leave alone: the established dry-goods lanes don't need any of this scrutiny. Four years of history, over 300,000 deliveries, a 91% match rate that keeps holding steady. Full autonomy there has actually been earned, not just assumed on day one.
The lesson: a client's excitement about "day one AI" is not evidence a lane is ready for it. We let a sales pitch set an engineering bar, and it took a near-empty cold-chain buffer to teach us the difference between a workflow that sounds like a good candidate and one that actually is.
Now here is the same thing as a story
The short version above is what you'd say out loud in a room. Read this one for the fifty-three minutes on a Tuesday that decided whether Sarsen kept a brand-new hospital contract at all.
Ewart has driven refrigerated freight for fourteen years. He knows which back roads flood first, which loading docks close early on a Friday, and which bridges have height limits nobody bothered to repaint the sign for. He didn't need Drayline to do his job well. He used it because it made a good morning better.
For the pharmacy lane's first two weeks, that held up fine. The highway legs were the same highway legs Drayline already knew from a hundred other lanes, and the model's suggestions matched what Ewart would have picked himself, almost every time. He checked the route against his own knowledge before every single run anyway, because the lane was new and he hadn't decided yet what to trust.
By week three, he only glanced at the ETA. By week five, he just followed the blue line on the dash without a second look, because on this lane, in its short life so far, it had never once been wrong.
Nobody told him to stop checking. A route that's never once been wrong makes that decision for you.
Then came a Tuesday. Ten days earlier, the county had closed a lane for bridge repair, half a mile past the turn Ewart usually took. Drayline had never driven that stretch enough times to have learned the closure existed, so it routed him straight through it, with the same flat confidence it uses on a road it has run three hundred thousand times.
He caught the barrier at the last second and doubled back the long way. Fifty-three minutes gone, on a lane where the whole insulated window was only 240 minutes wide to begin with.
We didn't take fifty-three minutes from Ewart's morning. We nearly took a hospital's entire vaccine order.
He wasn't careless. He did the sensible thing every single week: a route that's never once been wrong earns less double-checking over time, not more. That's not a flaw in him. It's what anyone does with a tool that never once hedges.
The dock supervisor's call reached Orlagh before lunch. She didn't call it a model bug in the retro, because it wasn't one. Months earlier, in the sales meeting that closed the contract, "Drayline routes it from day one" had been the exact line that won the deal, and nobody in that room had asked whether a brand-new lane had anything like the data an established one did.
So instead, she pulled every lane Sarsen had signed in the past year and ran each one through a real check for the first time: a data floor of at least 500 completed deliveries and at least 90 days of local history before any lane left advisory mode, and a driver confirmation step for every route below that line. Under the new default, the pharmacy lane would have flagged that Tuesday's route as low-confidence before Ewart ever left the yard. He'd have taken his own known detour, the one he'd been driving for fourteen years, and the buffer would have stayed above 140 minutes the whole trip instead of dropping to 22.
One design let a confident wrong turn walk straight into a cold-chain deadline. The other would have caught it before the truck ever left the dock.
Here's what I'd tell myself, the day we sold that contract as day-one AI: a signature is not evidence a lane is ready. I mistook a sales promise for proof, and Ewart very nearly paid for it with someone else's medicine.
The five letters behind Drayline's new-lane rule
PICK would fit if this were one lane, one yes-or-no call. Here Orlagh has dozens of candidate lanes competing for the same small data and engineering team, and the job is ranking all of them by what actually makes a workflow ready. That's ORDER's job.
Five checks a lane clears, in order, before it gets full autonomy.
OOutcome. What a good candidate check is actually protecting.
Not a sales team's promise, and not a tidy-looking product roadmap. Sarsen's limited data and engineering hours, going to workflows that will actually hold up once shipped, not to whichever lane landed the biggest contract that month. Every other letter exists to serve this one thing.
Name the outcome before ranking anything. Skip this and every characteristic below is just personal taste dressed up as a checklist.
RReversibility. Which mistake is hardest to walk back.
A lane where a wrong answer is severe and hard to catch, like a cold-chain shipment on a tight window, is far harder to recover from than a lane where a wrong answer is cheap and shows up the same day. That's exactly why the two highest-weighted characteristics below are the ones that keep a mistake catchable and keep it cheap.
This is why order matters, not preference. A dry-goods delay costs an apology. A missed cold-chain window costs a client and, worse, a shipment nobody can replace in time.
Reversibility isn't a reason to avoid the harder call. It's a reason to check the data before making one.
DDependency. What has to already be true before a lane is even a candidate.
Enough real route, traffic, and delivery-outcome history to learn the lane's actual pattern, not just what the general model assumes from everywhere else. The pharmacy lane had 260 deliveries across six weeks. The floor set afterward: at least 500 real runs and at least 90 days, both, before a lane leaves advisory mode. No amount of engineering skill closes that gap. Only more real, checked history does.
This is the step a normal feature request doesn't have. A new report or a new screen doesn't need a data floor behind it before anyone can build it. A route model does.
Every box after the second one only means something once the second one is actually done.
EEvidence. What's cheap to check before ranking a lane seriously.
A quick look at real volume, how often the decision actually gets made, and current handling cost, how long and how costly it is to do the lane by hand today. The pharmacy lane runs three trucks a day, and a dispatcher was already routing it manually in about 20 minutes each morning. That's not urgent, whatever the client wanted marketed as day-one AI.
This is the step that stops a small, exciting lane from jumping a large, quiet one. It costs an afternoon of counting, not a sprint of building.
The path from a signed contract to full autonomy, with a real floor to clear along the way.
RRank. The five characteristics, stated plainly and weighted in this order.
One, enough real, representative examples to learn from, the load-bearing check, since nothing below matters if this one fails. Two, a fast, cheap way to tell if an output was right, because slow feedback lets a bad pattern run for weeks unnoticed. Three, a clear, stated line for "good enough," since without one nobody can build an eval or say when the model has earned trust. Four, room for it to be wrong sometimes without real harm, which is why the same 58% match rate is fine on a dry-goods lane and dangerous on a cold-chain one. Five, enough repeats to be worth the build at all, which matters for payback speed but never overrides the first four.
If this rank would look identical with a different Outcome in step O, it was picked by gut. Swap the outcome to "win the most new logos this quarter" and volume jumps to the top, which is exactly why naming Outcome first matters.
One alternative is worth naming and rejecting directly: pausing all new-lane automation company-wide for a full quarter, which is roughly what Sarsen's ops team pushed for right after the incident. It lost, in the end, to advisory mode, since a blanket pause would have thrown away a routing engine that already works well on established lanes, in exchange for safety that a lane-by-lane data floor delivers just as well. The AI-specific failure worth naming is a model with too little local history answering with the exact same flat confidence it uses everywhere else, a kind of silent, unflagged guessing that has no real equivalent in a plain feature. The guardrail is the 500-run, 90-day floor plus a mandatory driver confirmation step below it. And the trade-off is accepted on purpose: advisory mode costs Ewart a few extra minutes of route review each morning on thin-data lanes, in exchange for never again coming within 22 minutes of losing a shipment that couldn't wait.
And if you want to be sure it really works, try it somewhere else
Same five checks, a crop-disease app instead of a freight router, and the honest answer this time is to say not yet.
Tillmark Agronomics runs LeafScout, a tool that reads a phone photo of a leaf and tells a farm advisor what's wrong with the plant. Field lead Caoimhe Corvo gets requests the way Orlagh does: every account manager has a favorite idea. One boutique client, a small dragon fruit grower, wants LeafScout to identify their crop's diseases in time for a trade show demo next month.
The request in the top-left corner is the one worth declining, not the one worth rushing.
Same steps, mapped onto Tillmark. Outcome: protect the small data science team's time for crops that can actually be learned well, not for whichever client is loudest before a trade show. Reversibility: a public demo announcement is far harder to walk back than an internal pilot with no date attached. Dependency: LeafScout's corn and soybean models run on 1.2 million labeled leaf photos across six growing seasons. Dragon fruit has about 900 tagged photos from 40 farms and one season, nowhere near that floor. Evidence: dragon fruit accounts for a tiny slice of total volume, and there's no dedicated extension network for it, so confirming a diagnosis can take a full season instead of the usual three days. Rank: this request lands in the "not yet" tier. The demo gets turned down, and photos route to a plant pathologist on retainer by email until real, checked data builds up.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to Rank, name the tiers and the floor, everything else is support.
Cost: there's no time to check the real photo count before the meeting. Say so plainly, and use whatever's cheap and real, a quick count from wherever the tagged photos live, instead of assuming the data is probably fine.
The model got better, for real: a newer vision model turns out to need far fewer labeled photos per crop to hit a usable bar. Say that too, and move the request up a tier, from real evidence, not from hope.
Where people run it wrong.
They skip the Dependency check for anything promised to a VIP client before a public date.
They mistake a good-looking demo for real feedback, when nobody has checked it against what actually happened in the field.
They let raw volume stand in for readiness, when a workflow can happen constantly and still have almost no real signal behind it.
How to use it live. Before saying yes to anything, ask out loud: "do we actually have enough real, checked examples of this, or are we hoping the model figures it out?" If the honest answer is "we don't know," that's the whole Dependency check, and it's reason enough to hold the request until someone does know.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits ranking which AI workflow gets built or automated next?
Tap to flip
ANSWER
ORDER: outcome, reversibility, dependency, evidence, rank. Built for ranking real candidates by what's hardest to undo, and for checking whether a workflow is even ready to be ranked at all.
2 · THE CAST
Who holds each role in this story, and where do they work?
Tap to flip
ANSWER
Orlagh Enright leads product for Drayline at Sarsen Freight Co. Ewart drives the new refrigerated pharmacy lane. Caoimhe Corvo runs LeafScout's field advisory team at Tillmark Agronomics, in Section 4.
3 · THE OUTCOME
What is a good candidate check actually protecting?
Tap to flip
ANSWER
Sarsen's limited data and engineering hours, going to workflows that will actually hold up once shipped, not to whichever lane won the biggest contract that month.
4 · REVERSIBILITY
Which kind of mistake is hardest to walk back on a new lane?
Tap to flip
ANSWER
One that's severe and hard to catch, like a cold-chain window slipping, not one that's cheap and shows up the same day, like a dry-goods delay.
5 · THE OLD DECISION
What decision would Orlagh take back?
Tap to flip
ANSWER
Making full autonomy the default setting for any newly signed lane, because that's what sales had already promised. She'd default new lanes into advisory mode instead, until they clear a real data floor.
6 · THE NUMBER
Fill in the blank: the pharmacy lane had ___ deliveries behind it. The established lanes had over ___.
Tap to flip
ANSWER
260 deliveries across six weeks, against 300,000 completed deliveries across four years on the established lanes.
7 · THE RANK
State the five characteristics, in weighted order.
Tap to flip
ANSWER
Enough real data to learn from, fast and cheap feedback on right versus wrong, a clear good-enough bar, room to be wrong without real harm, and enough repeats to be worth building.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs ORDER again on a different product. Which one, and who runs it?
Tap to flip
ANSWER
LeafScout, Tillmark Agronomics' crop-disease photo tool. Field lead Caoimhe Corvo uses the same method, and this time turns a dragon-fruit request down.
Check yourself Score: 0 / 0
Fill in the blank
1. The pharmacy lane had ___ completed deliveries behind it when it switched to full autonomy. The established lanes had over ___.
Show hint
Check flashcard 6 and the Dependency letter.
Show answer
260; 300,000. Sarsen's floor set afterward was 500 real runs and 90 days, both, before a lane leaves advisory mode.
Multiple choice
2. Why did Drayline route Ewart through a closed lane so confidently, with nothing on the screen suggesting it might be wrong?
A. A software bug in the GPS layer.
B. It had too little real local history on that lane to know the closure existed, and it showed the same confidence whether it was right or wrong.
C. Ewart ignored the suggested route.
D. The hospital asked for a shorter route.
Show hint
Check the Dependency letter and stage 5 of the walkthrough.
Show answer
B. The gap came from a real, checkable data shortage, not a bug or driver error, and nothing about the tool's own output signaled it was guessing.
True or false
3. True or false: in Orlagh's system, a lane with a big client and strong leadership support skips the data check.
True
False
Show hint
Check the Rank letter and the direct answer.
Show answer
False. The data floor applies to every lane regardless of who is asking, which is exactly the rule that was missing when the pharmacy lane shipped.
Short answer, name the rejected option
4. What did Sarsen's ops team push for right after the incident, and what did Orlagh choose instead?
Show hint
Check stage 6 of the walkthrough and the closing paragraph of the framework recap.
Show answer
Model answer: They pushed to pause all new-lane automation company-wide for a full quarter. Orlagh chose advisory mode instead: Drayline keeps suggesting routes, but a driver confirms them below the data floor, so the working engine on established lanes isn't thrown away.
Short answer, apply it yourself
5. Think of a repeated decision you make at work or at home. What's one sign it doesn't have enough real past examples behind it yet to trust a shortcut for it?
Show hint
Check the Dependency letter, what has to be true before a workflow is even ready to rank.
Show answer
Model answer: Any decision where you'd have to guess at "what usually happens" rather than point to real past cases you've actually checked against what happened next. If you can't say how often your shortcut has been right, you don't have enough examples yet, you have a hunch.
Short answer, work the number
6. If the cold-chain window had been 300 minutes instead of 240, would the same wrong turn still count as a near miss? Why or why not?
Show hint
Check the Reversibility letter and the buffer numbers in "Let's learn."
Show answer
Model answer: probably not. With a 300-minute window, a normal 165-minute run leaves 135 minutes of buffer. Add the 53-minute detour and the truck still arrives with 82 minutes to spare, comfortably inside the window. How severe a mistake is depends on how tight the tolerance already was, which is why Reversibility has to be checked case by case, not assumed from the workflow alone.
Before you close the answer
Why this works
Tests whether a candidate can tell a workflow that could theoretically use AI apart from one that's actually ready to trust, specifically by weighing real data and real feedback speed, not a generic list of "good project" traits dressed up in AI language.
Follow-up traps
"What if the client threatens to walk if you don't automate on day one?" Response: separate the promise from the automation. Offer the manual, dispatcher-led version immediately, with a real date for full autonomy tied to the data floor, not a wish tied to the sales calendar.
"Isn't a fixed number like 500 runs or 90 days just arbitrary?" Response: no, it's set from what the established lanes actually needed to hit a stated match-rate bar, not a round number picked to look rigorous, and it moves if real evidence says a newer model needs less.
If pressed
Advisory mode isn't a one-time gate. A lane only earns full autonomy once its suggested routes get followed without a manual override on at least 95% of runs for two straight weeks, not just once it crosses the 500-run, 90-day count. Hitting the floor on volume alone, with drivers still overriding constantly, means the lane stays in advisory mode regardless of how many days have passed.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.