ConceptFoundationalAI Opportunity & Model Strategy / When NOT to use AI / #1

List five conditions under which you should reject an AI solution outright.

The direct answer
Reject the AI draft outright the moment its words could change someone's pay, promotion, or job status, and that person never sees the draft before it counts against them. That one condition kills the proposal by itself, no matter how good the model tests. Four more conditions sit right behind it, and the priority list below names all five.
The five reject conditions, in order
  1. Reject it if the draft can shape pay, promotion, or a performance plan, and the employee never sees it first.Why: this is the one condition that alone makes a mistake unrecoverable. By the time anyone checks, the decision built on it is already made.
  2. Reject it if sounding complete means the model has to invent a detail your notes never gave it.Why: an audit of one HR platform's drafts found 3 in every 100 added a specific fact nobody wrote down. Those are the ones that cause the trouble later.
  3. Reject it if there is no cheap way to catch the mistake before it becomes the permanent record.Why: a clumsy sentence gets fixed on the read-back in thirty seconds. A wrong fact that reads smoothly gets fixed five months later, for real money and a grievance file.
  4. Reject it if the mistake rate is worse for one group than another and the draft does not say which kind of review you are reading.Why: a model trained on years of review language can repeat the same unfair pattern for the same kind of employee, and every draft looks the same on the screen.
  5. Reject it if what you are really automating is the minutes a manager spends thinking about the person, not the typing.Why: some of what looks like slow, manual output is actually the manager doing their job. Speed that up and you have cut the part that mattered.

How to actually say this in the room

Seven moves. This is a commit question wearing a list's clothes, so say the position before you count anything on your fingers.

1
Reframe what "reject it outright" means
Say it like this
"Before I list anything, I want to say what 'reject outright' means here. It doesn't mean never use AI in performance reviews. It means there are five specific situations where I would stop this feature before it ships, no matter how good the demo looked."
Why this works
Interviewers ask this to see whether you refuse everything or accept everything. Naming the narrow scope up front shows judgment instead of fear.
2
State your position, the one that kills it alone
Say it like this
"If one thing is true, I don't need the other four. If this draft could change someone's pay, their promotion, or whether they keep their job, and that person never sees it before it's used against them, I reject it. Full stop."
Why this works
PICK rewards commitment. Naming the single hardest-hitting condition first, before any reasoning, is what separates a position from a list of worries.
3
Name who eats each kind of mistake, and how fast
Say it like this
"Split the mistakes into two kinds. A clunky sentence gets caught the moment the manager reads the draft back, thirty seconds, fixed, forgotten. A wrong fact that reads smoothly doesn't get caught at all. It goes straight into the file."
Why this works
This is the impact step. An interviewer wants to hear you separate visible cost from hidden cost, not talk about "quality" as one number.
4
Put a number on the asymmetry
Say it like this
"At one company I looked at, an audit of 1,200 drafted reviews found that 3 in every 100 had a specific detail, a number, a habit, a missed deadline, that never appeared anywhere in the manager's own notes. Ninety-seven were fine. That 3 in 100 is small on a slide and enormous in someone's file."
Why this works
A percentage with nothing attached to it is a fact nobody remembers. A percentage attached to a personnel file is the sentence that gets repeated in the debrief.
5
Give the five conditions, in order
Say it like this
"Here are the five, checked in this order. One, it can shape pay or promotion and the person can't see it first. Two, it has to invent a detail your notes never gave it. Three, nobody can catch the mistake before it becomes the record. Four, the mistake rate favors one group over another and you can't tell which draft you're reading. Five, you're automating a conversation, not a chore."
Why this works
This is the direct payoff of the whole question, said as a checklist a hiring manager could use tomorrow, not a theory.
6
Say what you would still ship
Say it like this
"None of this kills the feature. I'd still draft the plain summary of a good quarter, the goal-tracking recap, anything a manager reads out loud and owns before it goes anywhere. That's most of the reviews written all year."
Why this works
Five reasons to say no can sound like fear. One clear place you'd still say yes is what makes the rest sound like judgment instead.
7
Close on what would change your mind
Say it like this
"I'd change my mind the moment two things are both true: the manager sees and signs every specific line before it's saved, and we can show the mistake rate is the same across every group of employees, not just on average. Until then, the specifics on anything touching pay stay human."
Why this works
A pick with no kill criteria sounds stubborn. This line is what makes the whole answer sound like a standard, not a fear.

One more thing before the walkthrough moves on: these five conditions apply to a piece of the product, not the whole thing. Most candidates hear "list conditions to reject AI" and answer as if the question were "when do you turn the feature off." Say which part of the workflow each condition actually touches, and you've answered a much harder, much better question.

Let's learn

The tool sits inside a company's HR software, right where a manager already goes twice a year to write performance reviews. The manager types a few rough notes about how someone did. The tool pulls in the goals that person set for themselves last cycle. Then it turns both into a paragraph the manager can read, fix, and send.

A manager's rough notes and an employee's stated goals flow into an AI-drafted paragraph, which the manager edits into the official record
One note, one goal, one paragraph that outlives the meeting

Before the tool, writing one review by hand took about 35 minutes. A manager with nine direct reports spent just over five hours on that, twice a year, stacked on top of a normal week.

With the tool, the same review takes about 9 minutes. The manager reads the draft, fixes a line, sends it. Twenty-six minutes back, per review, twice a year, for every manager in the company.

Here is the turn. The extra speed is not the problem. The problem is what kind of mistake a faster tool makes, and who ever finds out about it. A company using the tool ran a compliance check on 1,200 drafted reviews and found that 3 in every 100 contained a specific detail, a number, a habit, a missed deadline, that never appeared anywhere in the manager's notes or the employee's goals. The model needed something to say, so it said something plausible instead of nothing.

A clunky sentence gets caught the same minute someone reads it. A wrong fact that reads smoothly doesn't get caught at all.

Ninety-seven drafts out of a hundred were fine. That is a good number on a slide. It is the wrong number to build a decision on, because the three that weren't fine didn't fail loudly. They read like every other line in the review, got signed, and went into a file that follows a person to their next promotion meeting.

Knowledge spark: what a made-up detail actually looks like Not a wild claim. A small, ordinary-sounding one. "Consistently misses inventory deadlines." "Has raised this with the team before." Nothing dramatic, which is exactly why nobody stops to check it. The model isn't lying on purpose. It's filling a gap in short notes the same way it fills a gap in any other sentence, with the most likely-sounding words.

The choice I would take back. Early on, the team decided the draft should always come back as a full paragraph, even when the manager's notes were three words long. That felt like the kind choice. Nobody wants a blank box staring back at them. I would take it back. If the notes are thin, the tool should say so, not fill the gap with something that sounds true.

What I would leave alone The short check-in note a manager reads out loud in a five-minute meeting, where both people react to it on the spot. If the tool invents a detail there, the employee corrects it in the room, before it goes anywhere. Nothing about it becomes permanent. Leave that one alone. The same risk that is dangerous in an annual review costs nothing here.

The lesson. We built the tool to save time, and it does. But we measured the time it saved and never measured what kind of mistake it made instead. A faster wrong answer is not progress. It just moves the cost somewhere we weren't looking.

One company, 3,800 reviews a year
No AI
Full AI
The split
Manager time, all in
$66,500
$17,100
$31,920
Fixing contested reviews
$45,600
$273,600
$54,000
Total for the year
$112,100
$290,700
$85,920
Turning AI loose everywhere saves the most manager time and costs the most overall, because it also fabricates the most mistakes. The split, AI only where a mistake is caught in the room, costs less than doing nothing by hand at all.
What each policy costs a year
Manager time
Fixing contested reviews
No AI, all by hand
$112,100
Full AI everywhere
$290,700
The split (my pick)
$85,920
The middle bar looks like the "generous" option and is actually the most expensive one, because the hidden mistakes cost more than the time it saves. The bottom bar wins on both axes at once, which is what a real asymmetry looks like once you price it.

The quarter that changed what Fernhill would ship

You don't need this to answer the question. Read it if you want to feel why the five conditions are specific instead of vague.

Yolanda Pruitt can walk a grocery floor once and tell you which of her nine store managers is about to have a rough quarter, before the numbers say so. She's been a district manager at Halloway Grocers for eleven years, and every March and every September she sits down to write a review for each of them.

Fernhill's drafting tool arrived the spring before last, and for three cycles it was just good. Writing nine reviews used to eat most of a Thursday. Now she typed a few lines about each manager, pulled up their stated goals from the system, and had a draft back in a minute. She read it, tightened a sentence, signed it. Done with all nine before lunch.

She never worried the drafts were wrong in a way that mattered. If a sentence read stiff, she rewrote it. If a line felt generic, she deleted it. Those fixes took her seconds, and she didn't think about them again.

Two boxes of unequal weight: awkward phrasing caught in thirty seconds, versus a made-up detail not caught for five months and filed as the record
Same tool, two very different mistakes

Then, last September, Fernhill ran a compliance audit across every customer using the tool, comparing 1,200 drafted comments against each manager's own notes, word for word. Yolanda's district turned up in the sample. One review had a line in it: "has missed three inventory deadlines this quarter." Yolanda's notes said nothing about deadlines. She hadn't written it, checked it, or noticed it. It read exactly like every other sentence in the draft, so she signed it in March, and it sat in that store manager's file for six months before anyone looked twice.

He never even knew it was there until the audit surfaced it and HR had to pull the file. He hadn't asked to see his own review since the day he signed off on it, because nobody at Halloway had ever told him he could.

We didn't build a tool that makes mistakes. We built one that makes mistakes nobody is looking for.

Renata Souza had been the product manager on the drafting feature since before it launched. When the audit results landed on her desk, the header number was fine. Ninety-seven drafts in every hundred matched the notes exactly. Three didn't. She almost let it go at that.

Then she read what the three actually said. Every one was a specific claim, a habit, a number, a missed deadline, that the model had invented to make a short set of notes sound like a full paragraph. None of the invented lines were cruel. That was the part that worried her most. They were small, plausible, forgettable, exactly the kind of sentence nobody double-checks, and exactly the kind that turns up eighteen months later in a promotion committee's file.

She didn't ask engineering to lower the mistake rate. Three in a hundred was already low. She asked a different question: which three in a hundred can we actually afford. A generic sentence in a check-in costs nothing. A made-up one in a review that touches pay or promotion costs a person their record, and they never got a vote on it.

Fixing that one review, once HR found it, took about three weeks and roughly $2,400 in combined HR and legal time to investigate, correct the file, and answer the store manager's questions. On a review that never touches pay, the same kind of fix runs closer to $600, because nobody has to prove anything to anyone; they just cross it out. Multiply the expensive kind by even a small share of drafts across every customer Fernhill has, and the ninety-seven good ones stop being the number that matters.

So the decision Renata made wasn't to shrink the feature. It was to draw a line through it. Full drafting stays on the check-ins and goal recaps a manager reads out loud in the room, where both sides can correct it on the spot. On anything that feeds a pay or promotion decision, the manager writes every specific themselves; the tool only organizes the notes into a structure. Nine reviews a year for someone like Yolanda still take less time than they used to. They just don't take zero of her attention.

The part Renata would take back isn't the launch. It's the early call that the draft should always come back sounding finished. A short, honest "not enough here to draft from" would have cost a manager ten extra seconds. Instead it cost one store manager six months of not knowing what was sitting in his own file.

PICK, one letter at a time

This is a tradeoff question dressed up as a list, so PICK is the tool, not FLIPS. A "how would you measure this" question would reach for LEAD instead.

P, position. Reject the AI draft outright the moment it could shape someone's pay or promotion and they never see it first. Keep it everywhere the manager reads the draft out loud and the person can push back on the spot.
I, impact. A generic line in a check-in costs the manager a few seconds to fix, and the employee never even notices, because it's said in the room and moved past. A made-up specific in a pay-linked review costs the employee their own record, unseen, for however long it takes someone to check it, which at Halloway was six months.
C, cost asymmetry. The cheap side is a few seconds of editing, over and over, forever, and everyone can see it happen. The expensive side is rare, three in a hundred, but invisible until an audit or a dispute forces someone to look, and by then it isn't a sentence anymore, it's a personnel file. Spend your caution on the one nobody's watching.
K, kill criteria. Five conditions flip this from "draft it" to "reject it outright": it can move pay or promotion with no chance to see it first; it has to invent a detail your notes never gave it; nobody can catch the mistake before it becomes the record; the mistake rate is worse for one group and the draft doesn't say which; or you're automating the minutes a manager spends thinking about a person, not the typing. Any one of these is enough on its own.
Knowledge spark: what makes something a kill criterion A kill criterion is a fact you could actually go check, not a feeling. "The model seems risky" isn't one. "It invented a detail the source notes never gave it" is, because you can point to the exact sentence and the exact note it doesn't match.
The fabrication rate, quarter by quarter, against the line that matters
Drafts with an invented detail, measured each quarter
Kill line for anything that touches pay or promotion, 1 in 100
0% 3% 6% 1% kill line 5.0% 4.2% 3.6% 3.0% Q1 Q2 Q3 Q4
Four model updates in a row made the rate better. It's still three times over the line that matters for pay-linked reviews. Better is not the same question as safe enough, and the chart is what keeps those two questions from getting blurred together.

Try the same five conditions on a claims desk

An auto insurer uses a model to draft the letters an adjuster sends about a claim. Same shape of question: five conditions, different desk, different paper trail.

P. Reject the draft outright the moment it could state fault or a dollar settlement figure the claimant hasn't seen a person confirm. Keep it for routine status updates and document checklists.
I. A wrong line in a routine status update gets caught in one phone call, costs about five minutes. A liability statement that shouldn't have been written gets read into a lawsuit as the company's own words, months later, and by then it's evidence, not a typo.
C. Status-update mistakes are common and cheap, caught constantly, fixed on the spot. Liability misstatements are rare and can't be taken back once they're mailed. Spend the caution where the mistake is permanent.
K. Reject it outright if the draft could read as an admission of fault, if it sets a dollar figure nobody with authority approved, if the claimant would get it before an adjuster reviews every line, or if the model has to guess at a detail the file doesn't contain. Any one of those, and a person writes that letter by hand.

What I would leave alone, on the claims desk The internal note an adjuster jots for their own file, never sent to anyone. If the draft gets a small detail wrong there, the adjuster catches it the next time they open the claim, because it's their own working notes, not a letter that left the building.

Swap the trigger and it still runs

  • Speed: the draft comes back in two seconds instead of two minutes. Doesn't move the line. The five conditions are about what the sentence can do once it's sent, not how fast it arrived.
  • Cost: the tool triples in price. Also doesn't move the line, once a fabricated line is what's actually expensive, not the subscription.
  • The model gets better: the fabrication rate drops from 3 in 100 to 1 in 1,000. Move the line, don't erase it. Anything that still touches pay or promotion still gets a human's eyes, because the cost of the rare miss hasn't changed, only its odds.

Where people run it wrong

  • Treating the average pass rate as the whole answer, when the three misses in a hundred are the only ones that ever cost anything.
  • Banning AI everywhere the moment one condition applies anywhere in the product, instead of drawing the line at the specific decision it actually touches.
  • Writing "keep a human in the loop" as the fix, without saying which loop, which human, and what they're actually checking for.

If you are asked this cold

Say the reframe out loud before you list anything. "Give me a second, I want to separate what this could break from how often it breaks." That's true, it's already stage one, and it buys you the ten seconds you need to find five real conditions instead of five vague ones.

Flashcards (click a card to flip it)

1 · THE FRAMEWORK
Which framework fits this question, and what is its hardest step?
Tap to flip
ANSWER
PICK, for tradeoffs. The hardest step is C, the cost asymmetry: naming which mistake is cheap and visible and which one is hidden and expensive, then building the answer around the hidden one.
2 · THE PERSON
Who is this answer about, and what did she notice that a launch review usually misses?
Tap to flip
ANSWER
Renata Souza, product manager on the review-drafting feature at Fernhill HR. She read past the 97 percent pass rate to ask what the other three drafts in a hundred actually said.
3 · THE HABIT
What did Yolanda stop double-checking once the draft got fast?
Tap to flip
ANSWER
Whether a specific detail in the draft was actually in her own notes. The sentences read so smoothly that she stopped treating them as guesses.
4 · THE ASYMMETRY
Name the two kinds of mistake here and what each one costs.
Tap to flip
ANSWER
A clunky or generic line: caught on read-back, fixed in seconds, forgotten. A made-up specific: not caught for months, becomes part of an employee's permanent record, costs about $2,400 to investigate and correct once it's found.
5 · THE POSITION
State the pick in one sentence, the way you'd say it out loud.
Tap to flip
ANSWER
Reject the AI draft outright the moment it could shape someone's pay or promotion and that person never sees it before it counts against them. Keep it everywhere the draft gets read out loud and corrected on the spot.
6 · THE NUMBER
Fernhill's audit found ______ drafts in every 100 contained a detail the manager's notes never gave it.
Tap to flip
ANSWER
3. Ninety-seven were fine. That 3 in 100 is the number the whole decision turns on, because an average pass rate hides exactly the mistakes that matter most.
7 · THE KILL CRITERIA
Name the five conditions that mean you reject the draft outright.
Tap to flip
ANSWER
It can move pay or promotion with no chance to see it first. It has to invent a detail your notes never gave it. Nobody can catch the mistake before it becomes the record. The mistake rate is worse for one group and the draft doesn't say which. You're automating a conversation, not a chore.
8 · THE TRANSFER
Section 4 runs PICK again on a different product. Which one, and what's the reject line there?
Tap to flip
ANSWER
An auto insurer's claim-letter drafting tool. Reject outright any draft that states fault, sets a dollar figure, or reaches the claimant before an adjuster reviews it line by line.

Check yourself Score: 0 / 0

Fill in the blank
1. Fill in the blank: Fernhill's compliance audit found that ______ drafts in every 100 contained a detail the manager's own notes never gave it.
Show hint
It's the number the whole decision turns on, not the 97 that were fine.
Show answer
3. Three sounds small until you remember what kind of file it lands in. A pass rate of 97 percent hides the fact that the 3 percent are the only ones anyone will ever dispute.
Multiple choice
2. Which of these is one of the five conditions for rejecting the AI draft outright?
  • A. The model's confidence score drops below 90 percent.
  • B. It can shape someone's pay or promotion, and they never see the draft first.
  • C. A manager takes longer than 10 minutes to edit the draft.
  • D. The draft uses a word the employee might not like.
Show hint
Three of these are dials you'd tune. One is a condition you'd actually reject over.
Show answer
B. A, C, and D are all things you'd adjust and keep shipping. B is the one condition that alone makes the harm unrecoverable, which is what a kill criterion has to be.
True or false
3. True or false: rejecting the AI draft outright means Fernhill should turn the feature off for every manager.
  • True
  • False
Show hint
Look at what Renata actually kept running.
Show answer
False. She drew a line through the feature, not a line through the whole feature. Drafting stayed on check-ins and goal recaps, where the employee reads and corrects it in the room. It only came off reviews that touch pay or promotion.
Multiple choice
4. Why couldn't Halloway just add one more review step instead of pulling AI off the high-stakes reviews?
  • A. Adding review steps was against company policy.
  • B. The made-up lines read exactly like every other sentence, so a second reader would miss them too.
  • C. Managers refused to do extra work.
  • D. Fernhill charges extra for a review step.
Show hint
Ask whether the mistake actually looks different from a correct sentence.
Show answer
B. "Add a review step" only works if the mistake stands out to whoever's reviewing. A made-up detail that sounds plausible passes a second read as easily as it passed the first. The fix has to change who owns the specifics, not add another pair of eyes reading the same sentence.
Short answer
5. If the audit had found 12 fabricated drafts in every 100 instead of 3, would the split still beat turning AI off everywhere? Walk through it.
Show hint
Redo the dispute-cost side of the ledger at four times the rate, and compare it to what turning AI off entirely would cost in manager time.
Show answer
Yes, by more, not less. At 3 in 100, full AI everywhere already cost more than no AI at all, because of the dispute cost. At 12 in 100, full AI everywhere gets dramatically worse. The split barely moves, because it already keeps AI off the reviews where a fabricated line is expensive. A higher fabrication rate makes the case for the split stronger, not weaker. It's exactly the kind of number that should make you tighten the line, not loosen it.
Short answer, apply it yourself
6. Pick a tool you use yourself. Name one thing it does where a wrong output gets caught right away, and one where a wrong output could sit unnoticed for months. Which one deserves a human checking every specific?
Show hint
Look for the output that goes somewhere you never look back at.
Show answer
Model answer: "Take an expense app that auto-categorizes charges. A wrong category on a coffee run gets caught the moment I glance at the monthly summary, and costs nothing to fix. But if it silently tags a real business expense as personal, that mistake sits in a tax filing for a year until an audit finds it. The category field deserves a person checking every entry that touches taxes. The coffee runs don't." Any answer works if you can name the output nobody looks back at.
From U2xAI Academy

From answering questions to owning outcomes.

A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.

  • A live AI agent you actually shipped
  • A launch decision you can defend under pressure
  • An interview-ready portfolio, not more flashcards
Know more