ConceptIntermediateAI Opportunity & Model Strategy / Opportunity identification for AI / #18

What makes an internal AI opportunity more attractive than a customer-facing one for a first project?

FLIPS · picking Rackspine's first AI build, an internal tagging model versus an employee-facing chat bot

Rackspine helps IT departments track every laptop, monitor, and software license a company owns. Saorise Cundall owns its AI roadmap, and for two roadmap cycles running she pitched the same idea as Rackspine's first AI feature: a chat bot any employee could ask about their gear. The quieter option, a tagging model only the IT team would ever see, never made the slide. Then Rackspine tested the chat bot on its own 340 employees first.

The direct answer
Build the internal asset-tagging tool first, not the employee-facing chat bot. Its first wrong guess lands on an IT admin who already checks a review queue every day, not on an employee with no way to know the model got it wrong. Ship the flashier, customer-facing feature second, once the internal build has produced real accuracy numbers and a working way to catch a mistake before it reaches anyone.
Do this, in order
  1. Build the internal asset-tagging tool first, not whichever idea reads best in a customer email.Why: this is the actual reversal; skip it and "which AI project first" quietly turns into "which one gets a headline."
  2. Rank the two candidates by who stands between the model and its first wrong guess, not by how each one demos.Why: an admin working a review queue and an employee typing into a chat box face completely different odds of ever catching the same mistake.
  3. Let the internal build run long enough to produce a real labeled eval set and a real accuracy number before greenlighting anything customer-facing.Why: Rackspine's chat bot launch bar, 96 percent on a stratified holdout, only existed because six weeks of the internal tool's flagged corrections built it.
  4. Carry the internal tool's "loop someone in" habit over to the customer-facing feature instead of rebuilding it from scratch.Why: the same confidence line that flagged tags for Petrine is what later tells the chat bot when to say "not sure, ask IT" instead of a wrong sentence.
  5. Score AI project proposals separately from ordinary feature proposals on the roadmap.Why: "how visible will this be to a customer" is the right question for a normal feature and the wrong one for anything that's allowed to be confidently wrong sometimes.
  6. Skip this whole sequencing rule for a genuinely low-stakes AI feature, like a suggested nickname for an asset record.Why: nobody's laptop gets returned over a wrong nickname, so building internal-first there buys nothing but delay.

How to answer this, stage by stage

Nobody is grading whether you can rank two feature ideas nicely. They are grading whether you'll ask who's standing behind a model's first wrong guess, before you ask which idea will look better in an email.

1
Scope it to one product, one PM, one decision
Say it like this
"Let's ground this. Rackspine helps IT departments track laptops, monitors, and software licenses. Saorise Cundall owns its AI roadmap. Leadership wants Rackspine's first AI feature, and for two cycles running she's pitched the same idea: a chat bot any employee can ask about their gear."
Why this works
Naming the product, the person, and the two competing ideas stops the answer from staying a vague statement about picking a good first project.
2
Say your structure out loud
Say it like this
"I'll run this as FLIPS. Find the person whose scoping habit is at stake. Locate what she reaches for without thinking about it. Identify the flip, the exact question that changes. Pinpoint the old rule that only worked before anyone weighed the real risk. Show the replay with the internal tool built first."
Why this works
Two sentences of structure buy you a plan before a single detail arrives.
3
Reframe what's actually being tested
Say it like this
"This isn't really asking me to rank two feature ideas. It's asking whether I'll notice that 'first AI project' quietly gets scored by how good it'll look in an email, instead of by who's standing there the first time the model gets it wrong."
Why this works
Compresses the whole answer into one breath before a detail can bury it.
4
Give the one decision
Say it like this
"Here's what I'd actually do. Build the internal tagging tool first. It's boring, nobody outside IT will ever see it, and that's exactly the point: its first wrong guess lands on an admin already checking a queue, not a stranger with no reason to doubt a confident sentence."
Why this works
This is the direct answer, said plainly, before the story arrives to earn it.
5
Prove it with a compressed failure
Say it like this
"Here's what happens without it. Rackspine tests the chat bot on its own 340 employees first. A bulk import merges a loaner laptop's record into Fennrick Dwyre's. He asks if his laptop's still under warranty, and the bot tells him, flatly, that it's due back to IT. He boxes up his real work laptop that afternoon and drops it at the mailroom. He's without one for two days before anyone catches it."
Why this works
Four sentences carry a whole incident that a full retelling would take a page to earn.
6
Name the AI-specific detail you'd hold onto
Say it like this
"The real fix wasn't a smarter model, it was picking who stands behind its first mistake. Run the same kind of merged record through Rackspine's own asset-import queue instead, and Petrine Ashvold catches it in about ninety seconds, because it lands as a flagged tag in a queue she checks every morning, not a sentence with nothing behind it to make her doubt it."
Why this works
Shows the judgment is about who catches a probabilistic error, not about the model needing to be smarter.
7
Close on the one line
Say it like this
"So: 'which AI project should we build first' is never really a ranking question. It's a question about who's standing in front of the model's first wrong guess, and for a genuinely unproven first attempt, that has to be someone who can catch it fast, not a customer with no way to know."
Why this works
Leaves the interviewer with the actual decision, not just a well-told story about one near miss.

Let's learn

The review queue is a screen behind a login only the IT department ever sees. Rackspine is the software that fills it: it tracks a company's laptops, monitors, and software licenses, from the day a device gets ordered to the day it gets wiped and resold.

Before any AI touched it, an admin classified every new asset record by hand: what kind of device it is, which department owns it, when its warranty runs out. At Rackspine's own 340-person company, that fell to Petrine Ashvold every Monday, working through around 2,400 fresh records a month pulled from procurement invoices and a nightly network scan. Call it four hours a week, done by eye.

Two AI ideas sat on the roadmap to replace that Monday morning. One tagged the same records automatically, right in Petrine's queue, so she'd only open the ones the model flagged. The other put a chat box in the company's internal Slack: type a question about your gear, get an answer, no admin involved at all. Saorise Cundall, who owns the roadmap, pitched the chat box both times leadership asked what Rackspine's first AI feature should be. It was the one that would read well in a customer email.

Hand sketched two panel comparison titled The I step, in one picture. Left panel a gauge icon labeled The small move, caption a new AI idea lands, Saorise scopes it toward whatever will read best in a customer email, same as always. Right panel a question mark box icon in copper labeled The big snap, caption now the first question is who catches the model's first wrong guess, every time, before anything else.
Nothing about the mandate changed. What counted as a real first project changed completely.

Here's the turn. The two ideas don't fail the same way. Tested side by side on the same batch of records, both models get roughly the same slice wrong, about 1 in 9. But a wrong tag in Petrine's queue is a flagged row waiting for a decision. A wrong answer from the chat box is a sentence a stranger has no reason to doubt.

Share of the same wrong tags caught before they ever reach anyone outside IT
100% 50% 0 100% Internal review queue 0% Employee chat box, original design
Internal review queueEmployee chat box, no review step
Same model, same rough error rate. What decided the outcome was whether anything sat between the guess and the person acting on it.
We didn't ship Fennrick a wrong answer. We shipped him a stranger with no reason to doubt itself.

At its worst, that gap reaches a real person. During the internal alpha, a loaner laptop's record gets merged into Fennrick Dwyre's during a bulk import, a formatting mismatch on a serial number. He asks the chat box if his laptop's still under warranty. It tells him, plainly, that his device is past its refresh cycle and due back to IT. He boxes up his actual work laptop and drops it at the mailroom that afternoon. He's without one for two days, mid-project, because the bot never gave him a way to say that sounds wrong, and nobody was checking behind it.

The choice I would take back Rackspine's roadmap scores every idea the same way, on how visible it'll be to a customer. That rule is right for an ordinary feature, where the customer-facing version really is usually the one worth building first. It stops being right the moment the "feature" is a model that's allowed to be wrong sometimes, because now the real question isn't who sees it, it's who's standing there when it's wrong.
Knowledge spark: what's a confidence threshold? A cut-off point the model checks itself against before it answers. Below the line, it doesn't guess out loud, it flags the case for a person instead. Rackspine's tagger flagged anything under 92 percent confidence for Petrine to check by hand, instead of guessing anyway.

What I would leave alone: a small polish idea like an AI-suggested nickname for an asset record doesn't need any of this. Get it wrong and someone renames a laptop back to what they'd have called it anyway, costs nothing. The two-project sequencing earns its keep on ideas where a wrong guess actually costs someone something, not on every feature with a model behind it.

Hand sketched quadrant chart titled Where Rackspine's three AI ideas actually land. X axis how visible to a customer, from hidden to obvious. Y axis cost when the model is wrong, from cheap to expensive. Internal tagger sits low visibility low cost. Asset nickname sits low visibility very low cost. Ask Rackspine sits high visibility high cost.
Visible and costly-when-wrong are two different questions. Only one of them decided which idea got built first.

The lesson: visible was never the same question as safe to be wrong. Rackspine's roadmap had one box to score, how many people will see this, and for years that box was also, by accident, a decent stand-in for how good an idea was. It stopped being a stand-in for anything the day one of the ideas on it was allowed to guess.

Now here is the same thing as a story

The short version above is what you'd actually say in the room. Read this one when you want to feel exactly what a boxed-up laptop costs, not just hear the number.

Every quarterly roadmap review, Saorise Cundall walked in with the same kind of slide: whatever idea would look best in a customer email. She'd owned Rackspine's roadmap for three years, since before it had a single feature with a model in it, and she was good at the part of the job that mattered most in that room, turning a vague "do something with AI" into a slide leadership would actually fund.

Rackspine tracks a company's laptops, monitors, and software licenses. Its own 340-person office runs on its own product, the way a company that makes ladders keeps one in its supply closet. Two ideas had been sitting on the AI roadmap for a year. One was invisible: a model that tagged newly imported asset records, so Petrine Ashvold, the admin who did that by hand every Monday, would only open the handful the model wasn't sure about. The other already had a name: Ask Rackspine, a chat box in the company Slack that any employee could type a question into. Where's my laptop. Is my monitor still under warranty. Can I get a new headset.

The first time leadership asked what Rackspine's first AI feature should be, Saorise pitched Ask Rackspine. It demoed beautifully. Nobody in the room had ever seen the tagging model, and nobody was going to put "internal asset queue gets slightly faster" in a customer newsletter. Six months later, the same question came up again, this time with a real budget attached, and she pitched the same idea, polished further, still without ever running the two ideas side by side.

Hand sketched horizontal timeline titled Saorise's pitch, across two roadmap cycles. Four milestones: Cycle one, caption chatbot wins the room. Habit sets, caption no comparison run. Cycle two, caption same pitch again. The alpha, caption a near miss, this milestone emphasized in copper.
Nobody announced a change of plan. Two clean roadmap cycles were the whole decision, made without ever being made on purpose.

Then Rackspine did the thing it tells its own customers to do before shipping anything that guesses: it dogfooded the chat box on its own people first, three weeks before the real build. Fed it the company's own asset data. Turned it loose in the internal Slack for anyone at Rackspine to try.

A bulk import that week merged a returned loaner laptop's record into Fennrick Dwyre's, an engineer two desks from Petrine, because the loaner's serial number had been logged with a lowercase letter where his real device's record used an uppercase one, close enough for two different systems to treat them as the same machine. Fennrick asked the chat box, mostly out of curiosity, whether his laptop was still under warranty. The bot told him, in one flat, confident sentence, that his device was past its refresh cycle and scheduled for return.

He didn't doubt it. Why would he. He boxed up his actual work laptop that afternoon and dropped it at the mailroom for pickup, the way the bot's answer told him to. Nobody caught it for two days, because nothing in the chat box's design gave him a reason to ask anyone first, and nothing told anyone at Rackspine that a live, working laptop had just left the building on the strength of one wrong guess.

We didn't ship Fennrick a wrong answer. We shipped him a stranger with no reason to doubt itself.

The same week, the same kind of mistake happened on the other side of the building. A different bulk import merged two records the same way, this time inside the asset-tagging model's test queue. Petrine opened her Monday review, saw the flagged tag sitting where it always sits, needed about two seconds to see the department code didn't match the device history, and fixed it in about ninety seconds. She never knew it had almost been a Fennrick.

We considered the easy fix first: leave Ask Rackspine as the plan, just add a line under its answers, results may not always be accurate, check with IT. We killed that within a day. Fennrick didn't read fine print before he boxed up his laptop. He read one confident sentence from a tool built into the company's own Slack. A disclaimer doesn't change what that sentence looks like at 2pm on a Tuesday, it just moves the blame after the laptop's already gone.

Here's the decision I'd take back instead, and it isn't really Saorise's. Rackspine's roadmap scores every idea, the tagging model, the chat box, a redesigned invoice page, anything, on the same scale: how visible will this be to a customer. That rule made real sense for three years, because for three years, every idea on it was an ordinary feature, and an ordinary feature really is usually worth building where the most people will see it first. Nobody had ever needed a second scale, because nothing on the roadmap had ever been allowed to guess wrong on its own before.

Run the same kind of request again, six weeks later, with the tagging model built first instead. It launches quietly, catching about 1 in 9 records the same way it always would, flagged for Petrine, corrected in her queue, no incident, no headline. Each correction gets logged. By week six, the model's climbed from 89 percent accurate to 97, and Petrine's ninety-second checks have built something nobody planned for on purpose: 2,600 real, labeled examples of exactly the kind of mistake this model makes.

That number becomes Ask Rackspine's actual launch bar. Not a guess, not "it demoed well." Ninety-six percent on a stratified holdout pulled from real records, and a rule built straight out of Petrine's own habit: anything under the line doesn't answer, it says ask IT, and loops the question to a person instead of guessing. Ask Rackspine ships eleven weeks after the tagging model started, not the same week leadership first asked for it, and in its first quarter, not one employee gets told to return a laptop that's actually theirs.

What I'd tell the version of myself who wrote that first roadmap scale, back when nothing on it could be wrong on its own: the box that says "how many people will see this" isn't a bad question. It's just not the only one, the day something on the list is allowed to guess.

FLIPS, or the ninety seconds that decided which project came first

Not a trick to sound structured. It's the difference between an idea scoped for whoever will notice it fastest, and one scoped for whoever can actually catch it wrong.

Hand sketched numbered list titled FLIPS, five questions before an AI project gets picked. Five rows: F, find the person, Saorise, scoping Rackspine's first AI build. L, locate the habit, reach for whatever reads best in a customer email. I, identify the flip, who catches the model's first wrong guess, this row in copper. P, pinpoint the old decision, the roadmap scores every idea by visibility. S, show the replay, internal tool first, chatbot second, both proven.
Four setup and payoff letters, and one question that only had to be asked once a model was allowed to guess.
FFind the person. Whose habit is this?
Saorise Cundall, the PM who owns Rackspine's AI roadmap, the person whose pitch shapes which idea leadership actually funds first.
The flip belongs to whoever decides what counts as a real first AI project, not whoever eventually builds it.
LLocate the habit. What did she reach for without thinking?
Whenever leadership asked what Rackspine's first AI feature should be, she reached for whichever idea would read best in a customer email, the chat box, without ever laying the two ideas' real risk side by side.
That habit cost nothing while nothing on the roadmap could be wrong on its own. It stopped being free the moment one of the ideas was a model.
IIdentify the flip. What verb snaps?
Old setting: the first AI project gets scoped for the audience that will notice it, whichever idea reaches the most people fastest, employees, customers, anyone. New setting: the first AI project gets scoped for the smallest group that can actually catch it wrong, the review queue an admin already works, before it's trusted with anyone who can't. Nothing in between: there's no version where a genuinely unproven first model is a little bit customer-facing, its first mistake either lands on someone with a lever to catch it, or it doesn't.
This is the answer to the question in one line. An internal AI opportunity is more attractive for a first project because its first wrong guess is cheap to catch, not because it's smaller or safer to build.
PPinpoint the old decision. Which choice made sense before?
Rackspine's roadmap scores every idea, AI or not, by how visible it'll be to a customer. That was the right rule for three years of ordinary features, and nobody ever built a second scale for the day an idea on the list could be wrong on its own.
"Add a disclaimer to the chat box" would be a new dial bolted onto the same broken rule. Scoping the first attempt to the smallest population that can catch it is the reversal actually taken back.
SShow the replay. Same request, better ending?
A similar request lands again, six weeks after the near miss. This time the tagging model ships first, catches 1 in 9 records the way it always would, and Petrine's queue turns each correction into a labeled example. By week six it's 97 percent accurate with 2,600 real corrections behind it.
The replay ends in a count: Ask Rackspine ships eleven weeks later at a 96 percent holdout bar with a loop-in-IT rule built from Petrine's own habit, and its first quarter closes with zero laptops mistakenly returned.
Hand sketched full page metaphor titled What the roadmap assumed, and what was true. Left panel a gauge icon labeled DIAL, caption we assumed a good AI pitch just needs a bigger audience. Right panel a box icon in copper labeled SWITCH, caption either a person catches the model's mistake before it lands or nobody does, nothing between.
The whole answer, in one picture. Nobody designed a dial. Everybody got a switch, and for two roadmap cycles nobody had to throw it themselves.
Rackspine's tagging model, accuracy by week, the six weeks that built the launch bar
100% 92% 85% 89% 97%, wk 6 Wk 1 Wk 3 Wk 6
Weekly accuracy, tagging modelBecomes the launch bar
A chat box scoped for launch week one would have shipped on the 89 percent number. Six quiet weeks of Petrine's corrections bought eight more points before anyone outside IT saw it.
Hand sketched labeled parts diagram titled What Ask Rackspine's launch bar was actually built from. Central document icon labeled Ask Rackspine, with four radiating labels: 96 percent holdout bar. Confidence threshold. Loop in IT escape hatch. Six weeks of real corrections.
None of these four pieces existed before the tagging model ran. The replay isn't a better guess, it's a launch bar built from real corrections.

The AI-specific failure worth naming plainly is a merged-record misclassification, the kind of mistake a bulk import can cause on its own, reaching someone with no domain knowledge and no reason to doubt a confident sentence. The guardrail is the confidence threshold itself, tuned from a real eval set instead of a guess, under 92 percent inside the tagging queue, under 96 percent on Ask Rackspine's holdout, both routing to a person instead of answering. There's a real cost accepted on purpose here too: building the boring tool first delays the announceable one by about eleven weeks and spends real engineering time on something nobody outside IT will ever see, in exchange for a launch bar built from 2,600 real corrections instead of a demo that happened to go well.

And if you want to be sure it really works, try it somewhere else

Same five letters, a completely different flip family this time. Nobody's chasing a customer email here. A good result just quietly stops anyone from asking the same question twice.

Snoutbridge sells practice-management software to veterinary clinics: appointments, patient records, medication and supply tracking. Cormag Yusk owns its AI roadmap, and the first time his team built an AI feature, he got the sequencing right without anyone having to tell him twice: an internal triage helper, used only by the front-desk vet techs who type in a pet's symptoms, pre-fills a likely category for the vet to check before the appointment starts. Two clean quarters. Ninety-five percent accurate on the clinic's own structured intake fields. Not one incident.

Hand sketched flow diagram titled Cormag's greenlight process, the step that stopped happening. Four boxes connected by arrows: New AI idea. Weigh the risk, this box emphasized in copper. Who catches it question mark. Greenlight.
Same five questions, a completely different way the flip hides. This time the step that stopped happening was the one that used to run every single time.

The trap showed up on the second AI idea, not the first. Leadership wanted Ask Snoutbridge, a symptom checker pet owners could type into directly from the client portal, no vet tech in between. Cormag didn't reach for it out of habit the way Saorise did. He reached for it out of confidence: two good quarters told him the team knew how to build this now, so the same risk comparison that shaped the triage tool never got run a second time. Nobody decided to skip it. It just stopped feeling necessary.

A new hire sat in on the design review and asked one plain question: did we run the same accuracy check on this one that we ran on the triage tool? Nobody in the room had an answer.

What finally surfaced it The new hire wasn't trying to catch anyone out. She'd just joined from a team where that check was a standing item on every launch checklist, and its absence here read as an oversight rather than a rule. It took someone who hadn't lived through the two good quarters to notice the comparison was gone.

F · Cormag Yusk, who owns Snoutbridge's AI roadmap, the person whose sequencing habit shaped both AI projects.
L · Before greenlighting any AI feature, he used to run the same internal-first risk comparison every single time, regardless of how the last one went.
I · The over-trust flip, a different shape from Saorise's. Old setting: every new AI idea gets weighed against who catches its mistakes, no matter how well the last build went. New setting: two clean quarters are treated as proof the team's judgment itself has improved, so the comparison stops getting run at all, and the flashiest idea gets the green light on confidence alone. Nothing in between: either the comparison runs every time, or a good result quietly retires it.
P · The lesson from building the triage tool got written up as a retro note in a doc nobody opened again, instead of becoming a required gate every future AI proposal had to clear, because at the time, one success felt like proof enough.
S · The new hire's question forces the comparison to actually run, late. It surfaces something the triage tool never had to face: real pet-owner messages are messy, free text, nothing like the clinic's clean structured intake fields. Accuracy on the real language: 74 percent, against 95 on the tidy version. Ask Snoutbridge ships anyway, but in a reviewed-first mode for eight weeks, a vet tech checks anything under the confidence line before a pet owner sees it. By week eight the messy-language accuracy is up to 91 percent, and not one wrong triage guidance reaches a pet owner unreviewed in that window.

Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to it: the first AI project should go wherever a wrong guess is cheapest to catch, not wherever it's most exciting to announce.
Cost: no budget to build a full internal tool first. Even a two-week pilot with a dozen of your own staff beats shipping the customer-facing version untested, because a dozen people who can flag a mistake beat zero.
The model got better, for real: say the model's overall accuracy climbs on the next release. Doesn't matter, maybe it matters more. A model that's mostly excellent is exactly the one nobody thinks to keep checking on the population it was never actually tested against.

Where people run it wrong.
They treat "we already built one AI feature successfully" as proof the next one needs less scrutiny, instead of asking whether the new one's users look anything like the old one's.
They let a good demo, on any population, substitute for a real check against who'll actually use the thing.
They wait for a customer's bad day to force the comparison, when the whole point of running it first is that you don't need one to show up.

How to use it live. If an interviewer hands you two AI ideas and asks which one goes first, ask yourself one thing before answering: if this ships and gets it wrong on day one, who's standing there. If the honest answer is "a stranger with no way to know," that's the whole question, answered.

Flashcards (tap any card to flip it)

1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
A scope flip. Saorise's habit went from scoping the first AI project for the biggest audience it could reach, every employee at once, to scoping it for the smallest slice that could actually catch it wrong, the review queue an admin already works. The unit of who the model was allowed to be wrong in front of shrank.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Saorise Cundall, the PM who owns Rackspine's AI roadmap, a company that helps IT departments track laptops, monitors, and software licenses. Known for turning a vague "do something with AI" into a slide leadership would fund.
3 · THE HABIT
What did she stop doing because it worked?
Tap to flip
ANSWER
Whenever leadership asked what Rackspine's first AI feature should be, she pitched the employee-facing chat box, the idea that would read best in a customer email, without laying the two ideas' real risk side by side.
4 · THE FLIP, IN THIS STORY
What's the two-setting switch here?
Tap to flip
ANSWER
Old: the first AI project gets scoped for whoever will notice it fastest. New: it gets scoped for the smallest group that can actually catch it wrong before anyone else sees it. No version ships a little bit customer-facing.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Rackspine's roadmap scored every idea, AI or not, by how visible it'll be to a customer. Right for three years of ordinary features. Nobody built a second scale for the day an idea on it could guess wrong on its own.
6 · THE NUMBER
Fill in the blank: both ideas got about 1 in ___ records wrong in testing. The tagging model's error got caught in about ___ seconds. The chat bot's near miss cost Fennrick ___ days without a laptop.
Tap to flip
ANSWER
1 in 9; about 90 seconds; 2 days. Same rough error rate, completely different cost, depending on who was standing behind it.
7 · THE REPLAY
Same bad day, new design, what changes?
Tap to flip
ANSWER
The tagging model ships first, climbs from 89 to 97 percent accurate over six weeks, and turns Petrine's corrections into a real 2,600-example eval set. Ask Rackspine ships eleven weeks later at a 96 percent holdout bar with a loop-in-IT rule, and its first quarter closes with zero mistakenly returned laptops.
8 · CROSS-PRODUCT TRANSFER
Section 4 runs FLIPS again on a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Snoutbridge, a veterinary practice-management company. The over-trust flip: after two clean quarters building an internal triage tool, Cormag Yusk stopped re-running the risk comparison and greenlit a pet-owner-facing symptom checker on confidence alone, until a new hire's question forced the check to actually run.

Check yourself Score: 0 / 0

True or false
1. True or false: the tagging model and the chat box had very different error rates in testing, which is why one caused a near miss and the other didn't.
  • True
  • False
Show hint
Look at "Here's the turn" in Let's learn.
Show answer
False. Both models got roughly the same slice wrong, about 1 in 9. What differed was who stood behind each wrong guess: an admin in a review queue, or a stranger with no reason to doubt a confident sentence.
Fill in the blank
2. Rackspine's internal tagging model started at ___% accuracy and reached ___% by week six, the same six weeks that produced the eval set Ask Rackspine's launch bar was built from.
Show hint
Check the line chart in the framework recap.
Show answer
89%; 97%. Six weeks of Petrine's ninety-second corrections turned into 2,600 labeled examples, and that data became the chat bot's actual 96 percent holdout bar, not a guess.
Multiple choice
3. Why couldn't Rackspine have just added a disclaimer to the chat box instead of building the tagging tool first?
  • A. Legal blocked disclaimers by policy.
  • B. A disclaimer doesn't change what a confident sentence looks like to someone with no reason to doubt it; it only moves the blame after the laptop's already gone.
  • C. The Slack integration had no room for extra text.
  • D. Engineering refused to ship copy changes that quarter.
Show hint
Look at the paragraph where the disclaimer idea gets considered and killed.
Show answer
B. Fennrick read one flat, confident sentence, not fine print, before he boxed up his laptop. A caveat doesn't fix the underlying gap, it just relocates who gets blamed for it.
Short answer, name the reversal
4. What old decision does this answer take back, and why did it make sense when it was first made?
Show hint
Look at the key point box titled "The choice I would take back," in Let's learn.
Show answer
Model answer: Rackspine's roadmap scored every idea, AI or not, on how visible it would be to a customer. That made sense for three years, because every idea on the roadmap before this one was an ordinary feature, where the customer-facing version usually is the one worth building first. It stopped making sense once one of the ideas was a model allowed to be wrong on its own.
Short answer, apply it yourself
5. Think of two ways your own team could ship its first AI feature: one aimed at customers, one aimed at your own colleagues. What's the one number you'd want before picking either?
Show hint
Look for who stands between the model and its first wrong guess in each option.
Show answer
Model answer: How many of the population that option reaches has any way to flag a wrong answer before it's acted on. Not accuracy alone. A model that's right 89 percent of the time is a very different bet in front of people who can catch the other 11 percent than in front of people who can't.
Short answer, the number question
6. If the tagging model and the chat box had a 2 percent error rate instead of 1 in 9, would the same reversal still make sense? Why or why not?
Show hint
Think about what actually drives the decision: the error rate, or who's standing behind it.
Show answer
Model answer: Yes, for the same reason the answer never depended on the size of the error rate. Even a 2 percent miss still reaches someone with no way to catch it, versus someone who checks a queue every morning. The reversal is about who's standing there, not how often the model is wrong.
Before you close the answer
Why this works
Tests whether you'll treat "internal versus customer-facing" as a vibe about visibility, or as a real question about a model's reliability risk and who's standing behind its first mistake. Most candidates default to "start small" without ever naming what actually makes small safer.
Follow-up traps
"Isn't the internal option just boring, though? Won't leadership push back?" Response: keep the announceable idea as project two, not cancelled. It ships stronger once it's backed by a real eval bar instead of a guess, which is a better pitch to leadership too, not just a safer one.

"What if the idea genuinely only makes sense customer-facing, there's no internal population to test it on?" Response: then the guardrail moves inward instead of disappearing, a small labeled pilot with staff who can flag a mistake directly, before the general release, same principle at a smaller scale.
If pressed
Ask Rackspine's 92 percent internal flag line and its 96 percent external holdout bar weren't picked separately. The second number came directly from the first six weeks of Petrine's real corrections, not a fresh guess dressed up as a launch bar.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more