What makes an internal AI opportunity more attractive than a customer-facing one for a first project?
Rackspine helps IT departments track every laptop, monitor, and software license a company owns. Saorise Cundall owns its AI roadmap, and for two roadmap cycles running she pitched the same idea as Rackspine's first AI feature: a chat bot any employee could ask about their gear. The quieter option, a tagging model only the IT team would ever see, never made the slide. Then Rackspine tested the chat bot on its own 340 employees first.
- Build the internal asset-tagging tool first, not whichever idea reads best in a customer email.Why: this is the actual reversal; skip it and "which AI project first" quietly turns into "which one gets a headline."
- Rank the two candidates by who stands between the model and its first wrong guess, not by how each one demos.Why: an admin working a review queue and an employee typing into a chat box face completely different odds of ever catching the same mistake.
- Let the internal build run long enough to produce a real labeled eval set and a real accuracy number before greenlighting anything customer-facing.Why: Rackspine's chat bot launch bar, 96 percent on a stratified holdout, only existed because six weeks of the internal tool's flagged corrections built it.
- Carry the internal tool's "loop someone in" habit over to the customer-facing feature instead of rebuilding it from scratch.Why: the same confidence line that flagged tags for Petrine is what later tells the chat bot when to say "not sure, ask IT" instead of a wrong sentence.
- Score AI project proposals separately from ordinary feature proposals on the roadmap.Why: "how visible will this be to a customer" is the right question for a normal feature and the wrong one for anything that's allowed to be confidently wrong sometimes.
- Skip this whole sequencing rule for a genuinely low-stakes AI feature, like a suggested nickname for an asset record.Why: nobody's laptop gets returned over a wrong nickname, so building internal-first there buys nothing but delay.
How to answer this, stage by stage
Nobody is grading whether you can rank two feature ideas nicely. They are grading whether you'll ask who's standing behind a model's first wrong guess, before you ask which idea will look better in an email.
Let's learn
The review queue is a screen behind a login only the IT department ever sees. Rackspine is the software that fills it: it tracks a company's laptops, monitors, and software licenses, from the day a device gets ordered to the day it gets wiped and resold.
Before any AI touched it, an admin classified every new asset record by hand: what kind of device it is, which department owns it, when its warranty runs out. At Rackspine's own 340-person company, that fell to Petrine Ashvold every Monday, working through around 2,400 fresh records a month pulled from procurement invoices and a nightly network scan. Call it four hours a week, done by eye.
Two AI ideas sat on the roadmap to replace that Monday morning. One tagged the same records automatically, right in Petrine's queue, so she'd only open the ones the model flagged. The other put a chat box in the company's internal Slack: type a question about your gear, get an answer, no admin involved at all. Saorise Cundall, who owns the roadmap, pitched the chat box both times leadership asked what Rackspine's first AI feature should be. It was the one that would read well in a customer email.
Here's the turn. The two ideas don't fail the same way. Tested side by side on the same batch of records, both models get roughly the same slice wrong, about 1 in 9. But a wrong tag in Petrine's queue is a flagged row waiting for a decision. A wrong answer from the chat box is a sentence a stranger has no reason to doubt.
At its worst, that gap reaches a real person. During the internal alpha, a loaner laptop's record gets merged into Fennrick Dwyre's during a bulk import, a formatting mismatch on a serial number. He asks the chat box if his laptop's still under warranty. It tells him, plainly, that his device is past its refresh cycle and due back to IT. He boxes up his actual work laptop and drops it at the mailroom that afternoon. He's without one for two days, mid-project, because the bot never gave him a way to say that sounds wrong, and nobody was checking behind it.
What I would leave alone: a small polish idea like an AI-suggested nickname for an asset record doesn't need any of this. Get it wrong and someone renames a laptop back to what they'd have called it anyway, costs nothing. The two-project sequencing earns its keep on ideas where a wrong guess actually costs someone something, not on every feature with a model behind it.
The lesson: visible was never the same question as safe to be wrong. Rackspine's roadmap had one box to score, how many people will see this, and for years that box was also, by accident, a decent stand-in for how good an idea was. It stopped being a stand-in for anything the day one of the ideas on it was allowed to guess.
Now here is the same thing as a story
The short version above is what you'd actually say in the room. Read this one when you want to feel exactly what a boxed-up laptop costs, not just hear the number.
Every quarterly roadmap review, Saorise Cundall walked in with the same kind of slide: whatever idea would look best in a customer email. She'd owned Rackspine's roadmap for three years, since before it had a single feature with a model in it, and she was good at the part of the job that mattered most in that room, turning a vague "do something with AI" into a slide leadership would actually fund.
Rackspine tracks a company's laptops, monitors, and software licenses. Its own 340-person office runs on its own product, the way a company that makes ladders keeps one in its supply closet. Two ideas had been sitting on the AI roadmap for a year. One was invisible: a model that tagged newly imported asset records, so Petrine Ashvold, the admin who did that by hand every Monday, would only open the handful the model wasn't sure about. The other already had a name: Ask Rackspine, a chat box in the company Slack that any employee could type a question into. Where's my laptop. Is my monitor still under warranty. Can I get a new headset.
The first time leadership asked what Rackspine's first AI feature should be, Saorise pitched Ask Rackspine. It demoed beautifully. Nobody in the room had ever seen the tagging model, and nobody was going to put "internal asset queue gets slightly faster" in a customer newsletter. Six months later, the same question came up again, this time with a real budget attached, and she pitched the same idea, polished further, still without ever running the two ideas side by side.
Then Rackspine did the thing it tells its own customers to do before shipping anything that guesses: it dogfooded the chat box on its own people first, three weeks before the real build. Fed it the company's own asset data. Turned it loose in the internal Slack for anyone at Rackspine to try.
A bulk import that week merged a returned loaner laptop's record into Fennrick Dwyre's, an engineer two desks from Petrine, because the loaner's serial number had been logged with a lowercase letter where his real device's record used an uppercase one, close enough for two different systems to treat them as the same machine. Fennrick asked the chat box, mostly out of curiosity, whether his laptop was still under warranty. The bot told him, in one flat, confident sentence, that his device was past its refresh cycle and scheduled for return.
He didn't doubt it. Why would he. He boxed up his actual work laptop that afternoon and dropped it at the mailroom for pickup, the way the bot's answer told him to. Nobody caught it for two days, because nothing in the chat box's design gave him a reason to ask anyone first, and nothing told anyone at Rackspine that a live, working laptop had just left the building on the strength of one wrong guess.
The same week, the same kind of mistake happened on the other side of the building. A different bulk import merged two records the same way, this time inside the asset-tagging model's test queue. Petrine opened her Monday review, saw the flagged tag sitting where it always sits, needed about two seconds to see the department code didn't match the device history, and fixed it in about ninety seconds. She never knew it had almost been a Fennrick.
We considered the easy fix first: leave Ask Rackspine as the plan, just add a line under its answers, results may not always be accurate, check with IT. We killed that within a day. Fennrick didn't read fine print before he boxed up his laptop. He read one confident sentence from a tool built into the company's own Slack. A disclaimer doesn't change what that sentence looks like at 2pm on a Tuesday, it just moves the blame after the laptop's already gone.
Here's the decision I'd take back instead, and it isn't really Saorise's. Rackspine's roadmap scores every idea, the tagging model, the chat box, a redesigned invoice page, anything, on the same scale: how visible will this be to a customer. That rule made real sense for three years, because for three years, every idea on it was an ordinary feature, and an ordinary feature really is usually worth building where the most people will see it first. Nobody had ever needed a second scale, because nothing on the roadmap had ever been allowed to guess wrong on its own before.
Run the same kind of request again, six weeks later, with the tagging model built first instead. It launches quietly, catching about 1 in 9 records the same way it always would, flagged for Petrine, corrected in her queue, no incident, no headline. Each correction gets logged. By week six, the model's climbed from 89 percent accurate to 97, and Petrine's ninety-second checks have built something nobody planned for on purpose: 2,600 real, labeled examples of exactly the kind of mistake this model makes.
That number becomes Ask Rackspine's actual launch bar. Not a guess, not "it demoed well." Ninety-six percent on a stratified holdout pulled from real records, and a rule built straight out of Petrine's own habit: anything under the line doesn't answer, it says ask IT, and loops the question to a person instead of guessing. Ask Rackspine ships eleven weeks after the tagging model started, not the same week leadership first asked for it, and in its first quarter, not one employee gets told to return a laptop that's actually theirs.
What I'd tell the version of myself who wrote that first roadmap scale, back when nothing on it could be wrong on its own: the box that says "how many people will see this" isn't a bad question. It's just not the only one, the day something on the list is allowed to guess.
FLIPS, or the ninety seconds that decided which project came first
Not a trick to sound structured. It's the difference between an idea scoped for whoever will notice it fastest, and one scoped for whoever can actually catch it wrong.
The AI-specific failure worth naming plainly is a merged-record misclassification, the kind of mistake a bulk import can cause on its own, reaching someone with no domain knowledge and no reason to doubt a confident sentence. The guardrail is the confidence threshold itself, tuned from a real eval set instead of a guess, under 92 percent inside the tagging queue, under 96 percent on Ask Rackspine's holdout, both routing to a person instead of answering. There's a real cost accepted on purpose here too: building the boring tool first delays the announceable one by about eleven weeks and spends real engineering time on something nobody outside IT will ever see, in exchange for a launch bar built from 2,600 real corrections instead of a demo that happened to go well.
And if you want to be sure it really works, try it somewhere else
Same five letters, a completely different flip family this time. Nobody's chasing a customer email here. A good result just quietly stops anyone from asking the same question twice.
Snoutbridge sells practice-management software to veterinary clinics: appointments, patient records, medication and supply tracking. Cormag Yusk owns its AI roadmap, and the first time his team built an AI feature, he got the sequencing right without anyone having to tell him twice: an internal triage helper, used only by the front-desk vet techs who type in a pet's symptoms, pre-fills a likely category for the vet to check before the appointment starts. Two clean quarters. Ninety-five percent accurate on the clinic's own structured intake fields. Not one incident.
The trap showed up on the second AI idea, not the first. Leadership wanted Ask Snoutbridge, a symptom checker pet owners could type into directly from the client portal, no vet tech in between. Cormag didn't reach for it out of habit the way Saorise did. He reached for it out of confidence: two good quarters told him the team knew how to build this now, so the same risk comparison that shaped the triage tool never got run a second time. Nobody decided to skip it. It just stopped feeling necessary.
A new hire sat in on the design review and asked one plain question: did we run the same accuracy check on this one that we ran on the triage tool? Nobody in the room had an answer.
F · Cormag Yusk, who owns Snoutbridge's AI roadmap, the person whose sequencing habit shaped both AI projects.
L · Before greenlighting any AI feature, he used to run the same internal-first risk comparison every single time, regardless of how the last one went.
I · The over-trust flip, a different shape from Saorise's. Old setting: every new AI idea gets weighed against who catches its mistakes, no matter how well the last build went. New setting: two clean quarters are treated as proof the team's judgment itself has improved, so the comparison stops getting run at all, and the flashiest idea gets the green light on confidence alone. Nothing in between: either the comparison runs every time, or a good result quietly retires it.
P · The lesson from building the triage tool got written up as a retro note in a doc nobody opened again, instead of becoming a required gate every future AI proposal had to clear, because at the time, one success felt like proof enough.
S · The new hire's question forces the comparison to actually run, late. It surfaces something the triage tool never had to face: real pet-owner messages are messy, free text, nothing like the clinic's clean structured intake fields. Accuracy on the real language: 74 percent, against 95 on the tidy version. Ask Snoutbridge ships anyway, but in a reviewed-first mode for eight weeks, a vet tech checks anything under the confidence line before a pet owner sees it. By week eight the messy-language accuracy is up to 91 percent, and not one wrong triage guidance reaches a pet owner unreviewed in that window.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: the first AI project should go wherever a wrong guess is cheapest to catch, not wherever it's most exciting to announce.
Cost: no budget to build a full internal tool first. Even a two-week pilot with a dozen of your own staff beats shipping the customer-facing version untested, because a dozen people who can flag a mistake beat zero.
The model got better, for real: say the model's overall accuracy climbs on the next release. Doesn't matter, maybe it matters more. A model that's mostly excellent is exactly the one nobody thinks to keep checking on the population it was never actually tested against.
Where people run it wrong.
They treat "we already built one AI feature successfully" as proof the next one needs less scrutiny, instead of asking whether the new one's users look anything like the old one's.
They let a good demo, on any population, substitute for a real check against who'll actually use the thing.
They wait for a customer's bad day to force the comparison, when the whole point of running it first is that you don't need one to show up.
How to use it live. If an interviewer hands you two AI ideas and asks which one goes first, ask yourself one thing before answering: if this ships and gets it wrong on day one, who's standing there. If the honest answer is "a stranger with no way to know," that's the whole question, answered.
Flashcards (tap any card to flip it)
Check yourself Score: 0 / 0
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
Show hint
Show answer
"What if the idea genuinely only makes sense customer-facing, there's no internal population to test it on?" Response: then the guardrail moves inward instead of disappearing, a small labeled pilot with staff who can flag a mistake directly, before the general release, same principle at a smaller scale.
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.
- A live AI agent you actually shipped
- A launch decision you can defend under pressure
- An interview-ready portfolio, not more flashcards
More on Opportunity identification for AI
- #1 What characteristics make a workflow a good candidate for AI? List five.
- #2 Describe a method for finding AI opportunities inside an existing product without starting from the technology.
- #3 How do you distinguish a problem AI solves from a problem AI merely touches?
- #4 Rank these by AI suitability and justify: expense approval, contract review, invoice matching, hiring decisions.
- #5 Explain why high-volume, low-stakes, tolerant-of-error tasks are the best first targets.
- #6 Your support team handles 8,000 tickets a month. Structure a discovery process to find the AI opportunity.