Artifact critiqueAdvancedAI Opportunity & Model Strategy / When NOT to use AI / #21
Write the one-page memo recommending against an AI initiative your CEO championed.
LEADthe AI forecaster lost to the spreadsheet it was supposed to replace, by three real points
Skyloom sells B2B software. Julienne Desroches, its CEO, championed QuotaSense, an AI model meant to replace the sales team's simple weighted-pipeline forecast. Anneke Vosloo is the AI PM who backtested it before the company-wide rollout, and has to write the memo saying it isn't ready.
The direct answer
Kill QuotaSense's rollout, not the initiative outright. Backtested against six real quarters of Skyloom's own sparse, seasonal deal data, it performed no better than the existing weighted-pipeline method already in use, worse, in fact, by three real points of forecast error. Recommend a smaller, cheaper next step instead of a full replacement nobody has evidence for yet.
Do this, in order
Back the recommendation with a real backtest against real historical data, not a demo.Why: this is the whole memo. Without it, "I have concerns" loses to a CEO's conviction every time.
Name the specific number the model needs to beat, and whether it did.Why: a vague "it's not ready yet" gives nobody a clear bar to check the recommendation against.
State a concrete alternative next step, not just a "don't."Why: a memo that only says no reads as blocking. One with a real next step reads as partnership.
Say plainly what's still promising about the original goal.Why: the CEO's instinct, better forecasts matter, was right. Only this specific model, on this specific data, wasn't ready.
Keep it to one page, and lead with the number, not the caveats.Why: a memo buried in hedges is as useless as one with no evidence at all, the recommendation has to be findable in ten seconds.
Offer to revisit the moment new evidence exists.Why: this is a decision based on today's evidence, not a permanent verdict on AI forecasting at Skyloom.
How to answer this, stage by stage
Nobody is scoring whether you can write a firm memo. They're scoring whether the memo would actually survive being read by the person who championed the thing it's about.
Stage 1
Scope it to one real initiative, not "pushing back on leadership" in general
Say it like this
"Let me give you a real case. A CEO champions QuotaSense, an AI sales forecaster meant to replace the team's spreadsheet method. I backtest it against six real quarters before rollout, and it comes out worse than what it's replacing. That's the memo I have to write."
Why this works
Grounds the answer in a checkable result instead of a general theory about disagreeing with executives.
Stage 2
Say your structure out loud before any content
Say it like this
"I'll run this as LEAD. Link, the real metric the memo has to hang on. Early signal, how this got caught before full commitment. Abuse, the two ways a memo like this usually fails. Decision, the actual memo, written out."
Why this works
Signals a repeatable method for writing hard feedback, not a one-off act of courage.
Stage 3
Reframe the question: this isn't about disagreeing with the CEO
Say it like this
"The CEO's underlying instinct, better forecasts help the business, is completely right. This memo isn't disagreeing with that. It's saying this specific model, tested against this specific data, isn't there yet, which is a narrower and much more defensible claim."
Why this works
This is where a strong answer separates from framing it as a personal or political conflict.
Stage 4
Give the one decision: the memo's real recommendation
Say it like this
"Here's what I'd actually write. Recommend pausing the full rollout, not the initiative. State the backtest result plainly: QuotaSense's forecast error was 34 percent against the existing method's 31 percent, on our own real data. Propose blending the model's signal into the existing pipeline as an input, not a replacement, while more data accumulates."
Why this works
This is the direct answer, stated as an actual memo recommendation, not a vague expression of concern.
Stage 5
Prove it with the compressed evidence
Say it like this
"This isn't a hunch. I backtested QuotaSense against six real quarters, roughly 340 closed deals, genuinely lumpy and seasonal data. It missed by 34 percent on average. The pipeline method the team already trusts missed by 31 percent, over the same quarters. The thing we're replacing is currently more accurate than its replacement."
Why this works
This is where the story lives, compressed to the one comparison that actually proves the recommendation.
Stage 6
Name the AI-specific reason it fell short, and the real cost of proceeding anyway
Say it like this
"The honest reason is data, not effort. Our sales cycles are lumpy and seasonal, with too few closed deals per quarter for a model to reliably learn a pattern the pipeline method already captures with a simpler, hand-tuned weighting. Rolling QuotaSense out now would cost real dollars, real trust with the sales team, and a forecast that's measurably worse than what they already have."
Why this works
This is the load bearing judgment. It wouldn't make sense to ask this about a feature with no model in it, since the failure is specifically about data sparsity a model needs and a simpler method doesn't.
Stage 7
Say what stays promising, then close on the concrete ask
Say it like this
"This doesn't mean AI forecasting is a dead end at Skyloom. It means we don't have enough real data yet, and rolling this out on today's evidence would hand the sales team a worse tool than what they trust now. My ask: pause the full rollout, run QuotaSense as a secondary input for two more quarters, and revisit with real numbers, not another demo."
Why this works
Closes with a concrete, bounded ask instead of an open-ended objection, and restates the direct answer in one breath.
Let's learn
QuotaSense is a model meant to predict how much revenue Skyloom will close each quarter, replacing a spreadsheet where sales ops manually weights each open deal by its stage's historical close rate.
Four steps, none of them exotic. The backtest just asked the model to predict something Skyloom already knew the real answer to.
Julienne championed QuotaSense after seeing a competitor's case study, and the instinct behind it was reasonable: a company Skyloom's size shouldn't still be running its forecast off a hand-tuned spreadsheet. Before rolling it out company-wide, Anneke's team ran the backtest that should happen before any forecasting tool ships.
Forecast error by quarter, QuotaSense versus the existing weighted pipeline
QuotaSense, AI modelExisting weighted pipeline
QuotaSense's error runs higher in every single quarter, and spikes hardest in Q5, a seasonally unusual quarter the simpler method handled more gracefully.
Averaged across all six quarters, QuotaSense's forecast error came out to 34 percent. The existing pipeline method, the one it was meant to replace, averaged 31 percent over the same period. The replacement was measurably worse than what it was replacing.
We weren't debating whether AI forecasting could work someday. We were looking at six real quarters where it already didn't, against a spreadsheet nobody thought was cutting edge.
Here's the turn: the problem was never that QuotaSense was poorly built. The turn is that Skyloom's real sales data, roughly 340 closed deals across three years, lumpy and seasonal, simply isn't enough for a model to reliably out-learn a method sales ops had already hand-tuned against the same real patterns for years.
Knowledge spark: why does a model need more data than a hand-tuned method?
A weighted-pipeline forecast encodes a person's accumulated judgment directly, stage X deals close 40 percent of the time, built from years of watching it happen. A model has to rediscover that same pattern statistically, from examples, and with only a few hundred real deals split across seasons, there often isn't enough signal for it to learn what a person already knows.
Cost to build and run, QuotaSense versus the existing pipeline method
Build plus ongoing costAlready built, already free
A real, recurring bill to run something less accurate than the free method already in place.
At its worst, shipping QuotaSense as planned would have cost Skyloom real money for a measurably worse forecast, plus real trust with a sales team that would have quickly noticed their new tool was wrong more often than the spreadsheet they'd used for years.
The choice I would take back
Scoping QuotaSense's rollout timeline before a real backtest against Skyloom's own historical data existed. That made sense when the case study that inspired it looked compelling. It stopped making sense the moment a backtest was possible and nobody had insisted on running one first.
What I would leave alone: Skyloom's existing weighted-pipeline method stays exactly as it is for now, it's simple, explainable, and currently more accurate than its proposed replacement. The goal was never protecting it out of habit, the backtest just happened to show it was still winning.
The lesson: a competitor's case study is evidence about their data, not yours. Before replacing a working method with an AI model, backtest the model against your own real history and let that number, not the pitch, make the call.
Now here is the same thing as a story
The short version above is what you'd say out loud in the room. Read this one for what it actually felt like to write a memo you knew your CEO wouldn't want to read.
Anneke Vosloo had been Skyloom's AI PM for two years, and she genuinely liked working with Julienne. Julienne moved fast, trusted her team, and rarely needed convincing on good ideas. QuotaSense was the first time Anneke had real evidence Julienne's instinct was ahead of the evidence.
The backtest wasn't dramatic to run. Anneke pulled six quarters of real deal data, closed and open, and asked QuotaSense to predict what sales ops had already lived through. Then she ran the exact same six quarters through the existing pipeline method, the one everyone already trusted.
Not a story about AI being wrong in general. A story about this specific model, on this specific data, this specific quarter.
The numbers came back plainer than she expected. 34 percent average error for QuotaSense. 31 percent for the spreadsheet. Not a close call dressed up as one, a real, checkable gap in the wrong direction.
She sat with it for a day before writing anything. The easy version of this memo buried the number in three paragraphs of context and hedging. The other easy version was blunt enough to read as "your idea doesn't work," with nothing offered in its place. Neither one would actually survive being read by someone who'd championed this publicly.
The actual writing challenge wasn't finding the courage to say no. It was finding the version that would actually get read and acted on.
She wrote the real version: the number first, the recommendation second, a concrete next step third, all inside one page. No apology for the finding, no attack on the idea's premise either.
The memo wasn't there to win an argument. It was there to make sure the sales team didn't get handed a worse tool than the one they already trusted.
Julienne read it in the hallway, not a scheduled meeting, which Anneke had half-dreaded. "Show me the backtest," was the first thing she said. Not defensive, just direct. Anneke had the six-quarter chart ready.
The part of the pitch that never made it into the original case study Julienne had seen.
Anneke never had a fixed rule for when a CEO's pet initiative should get killed outright versus redirected. She had a feeling with two settings: the evidence says the underlying goal is wrong, or it says only this specific bet is wrong. This was clearly the second. Better forecasting was still worth wanting. This model, on this data, right now, wasn't the way to get it.
The fork the memo actually ran, stated plainly instead of buried in hedges.
Back when the rollout timeline first got scoped, moving fast on a promising case study wasn't an unreasonable call. It stopped being reasonable the moment a real backtest was possible and nobody had insisted on running it before the launch date got set.
Here's the replay: same model, same real data, but with the backtest run before the timeline instead of after. QuotaSense ships as a secondary input alongside the pipeline method, not a replacement, while two more quarters of real data accumulate. No wasted rollout, no sales team trust lost to a tool that measurably underperformed what it replaced.
One version of this story ships a worse forecast to the whole sales org and finds out from complaints. The other spends one afternoon on a backtest and turns a public commitment into a smaller, honest next step instead.
What I'd tell myself, staring at that 34 percent number the night before writing the memo: a good instinct and a ready model are two different things, and the only way to tell them apart is to actually check.
LEAD, run on a forecast that lost to the spreadsheet it was meant to replaceNot a script for softening bad news until it's harmless. LEAD is what makes a hard recommendation survive being read by the person who won't want to hear it.
L
Link. What real, measurable metric does the memo actually hang on?
Forecast error, backtested against six real quarters: 34 percent for QuotaSense, 31 percent for the existing pipeline method. Not a feeling, a number either side can check.
A memo with no real metric behind it is an opinion wearing evidence's clothes.
E
Early signal. How did this get caught before full commitment?
A backtest run before the company-wide rollout, using real historical data Skyloom already had, rather than trusting the competitor's case study that inspired the initiative.
The earliest possible check, run at the cheapest possible moment, before real money or trust was spent.
A
Abuse. How do memos like this usually go wrong?
Too hedged, and the real recommendation gets lost in three paragraphs of caveats. Too blunt, and it reads as personal pushback with nothing offered in its place. Anneke's draft avoided both by leading with the number and closing with a concrete next step.
This is the direct answer's real craft challenge: not whether to say it, but how to write it so it actually gets acted on.
D
Decision. The actual memo, written out.
See the full one-page memo below. States the number first, the recommendation second, a concrete next step third, and what stays promising about the original goal, all inside one page.
This is the direct answer to the question, as the literal deliverable it's asking for.
Internal memo, one page
To:
Julienne Desroches, CEO
From:
Anneke Vosloo, AI Product Manager
Re:
QuotaSense rollout: recommend pause, not cancellation
The finding
I backtested QuotaSense against six real quarters of our own deal data (roughly 340 closed deals) before the planned company-wide rollout. Average forecast error: 34 percent. Our existing weighted-pipeline method, over the same six quarters, averaged 31 percent. The tool meant to replace our current forecast is currently less accurate than it, on our own historical numbers.
Why this happened
Our sales data is genuinely lumpy and seasonal, too few closed deals per quarter for a model to reliably out-learn a method sales ops has already hand-tuned against these same patterns for years. This is a data problem, not an execution problem, and it isn't one more engineering effort fixes quickly.
What I'm recommending
Pause the full company-wide rollout currently scheduled for next quarter.
Run QuotaSense as a secondary input alongside the existing pipeline for two more quarters, collecting more real data without risking forecast quality.
Revisit with an updated backtest once that data exists, and commit to a rollout only if it clears the existing method's real accuracy.
What stays true
Better forecasting is still worth pursuing, and I don't think this data gap is permanent. I'd rather bring you a model that beats our own numbers than one that looks impressive in someone else's case study.
Happy to walk through the backtest data whenever works for you this week.
The recap, one line per letter: link is the real backtested error rate the whole memo hangs on, early signal is running that backtest before the rollout instead of after, abuse is avoiding both the over-hedged and the over-blunt failure modes, and decision is the actual one-page memo, numbers first, a real next step, and the door left open.
And if you want to be sure it really works, try it somewhere elseSame four letters, a customer-support triage tool instead of a sales forecast. This time the early signal comes from a pilot team's real complaints, not a formal backtest.
Tobias Renquist is the AI PM at Freightloom, a logistics software company. Freightloom's CEO championed AutoTriage, a model meant to auto-route every support ticket to the right team without human triage. Mapped onto LEAD: link is the real metric, percent of tickets correctly routed on the first try, backtested against 90 days of real tickets: AutoTriage hit 71 percent, the existing human triager hit 89 percent. Early signal came from a two-week pilot with one support pod before full rollout, where misrouted tickets visibly piled up in the wrong queues. Abuse means avoiding a memo that either buries the 71-versus-89 gap in soft language, or one that dismisses the CEO's underlying goal, faster ticket routing, which was still a real problem worth solving. Decision is a memo recommending AutoTriage handle only the clearest, highest-confidence 40 percent of tickets automatically, routing the rest to the human triager, instead of replacing the role outright.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Skip straight to "the model backtested worse than what it replaces, here's the number, here's the smaller next step," and stop.
Cost: no time to run a full backtest before the memo is due. Say so honestly, and recommend delaying the rollout decision until one exists, rather than guessing at a verdict.
The model got better, for real: say a later backtest shows QuotaSense finally beating the pipeline method. Write the follow-up memo just as plainly, with the new number leading it, the same way the first one did.
Where people run it wrong.
They write the memo as a personal objection instead of anchoring it in a real, checkable number.
They soften the recommendation so much the actual "don't roll this out yet" gets lost in the caveats.
They frame it as killing the initiative entirely, instead of pausing one specific rollout decision on today's evidence.
How to use it live. The moment an interviewer asks you to push back on a leader's pet initiative, ask yourself first: what's the one real number this recommendation would stand or fall on, and do I actually have it. That question buys real thinking time, and it's usually exactly what turns an opinion into a memo someone can act on.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits writing a memo against a CEO-championed AI initiative?
Tap to flip
ANSWER
LEAD: link, early signal, abuse, decision. It grounds the recommendation in a real metric, catches the problem before full commitment, and avoids both over-hedging and over-bluntness.
2 · THE PEOPLE
Who is this answer about?
Tap to flip
ANSWER
Anneke Vosloo, the AI PM at Skyloom who backtests QuotaSense. Julienne Desroches is the CEO who championed it.
3 · THE EVIDENCE
What real evidence does the memo's recommendation stand on?
Tap to flip
ANSWER
A backtest against six real quarters of Skyloom's own deal data, showing QuotaSense's average forecast error (34 percent) was worse than the existing weighted-pipeline method's (31 percent).
4 · THE ASK
What's the one concrete thing the memo actually recommends?
Tap to flip
ANSWER
Pause the full rollout, run QuotaSense as a secondary input alongside the existing method for two more quarters, and revisit once an updated backtest shows it clears the existing method's real accuracy.
5 · THE OLD DECISION
What old decision would this answer take back?
Tap to flip
ANSWER
Scoping QuotaSense's rollout timeline before a real backtest against Skyloom's own data existed. Reasonable when a compelling case study inspired it. Wrong once a real backtest was possible and nobody insisted on running it first.
6 · THE NUMBER
Fill in the blank: QuotaSense's average forecast error was ___ percent, versus ___ percent for the existing pipeline method.
Tap to flip
ANSWER
34 percent, versus 31 percent. The model meant to replace the spreadsheet was measurably less accurate than it, on Skyloom's own six quarters of real data.
7 · THE CRAFT
What are the two ways a memo like this usually fails, and how did Anneke's avoid both?
Tap to flip
ANSWER
Too hedged, burying the real recommendation, or too blunt, reading as personal pushback. Anneke's led with the number, stated the recommendation plainly, and closed with a concrete next step and what still stayed promising.
8 · CROSS PRODUCT TRANSFER
Section 4 runs LEAD again on a different product. Which one, and what's the equivalent decision?
Tap to flip
ANSWER
Freightloom's AutoTriage support-ticket router. The equivalent decision recommends AutoTriage handle only the highest-confidence 40 percent of tickets automatically, routing the rest to the human triager instead of full replacement.
Check yourself Score: 0 / 0
Multiple choice
1. Why did QuotaSense underperform the existing weighted-pipeline method?
A. The engineering team built it incorrectly.
B. Skyloom's real sales data was too sparse and seasonal for the model to reliably learn a pattern the hand-tuned method already captured.
C. Julienne set an unrealistic rollout deadline.
D. The sales team refused to use it.
Show hint
Look at the knowledge spark in "Let's learn."
Show answer
B. A model has to statistically rediscover a pattern from examples. With only a few hundred real deals across seasons, there often isn't enough signal to beat a method built from years of accumulated human judgment.
True or false
2. True or false: the memo recommends cancelling QuotaSense entirely and abandoning AI forecasting at Skyloom.
True
False
Show hint
Look at the memo's "What I'm recommending" section.
Show answer
False. It recommends pausing the full rollout and running QuotaSense as a secondary input for two more quarters, revisiting once it clears the existing method's real accuracy.
Fill in the blank
3. Fill in the blank: the backtest used roughly ___ real closed deals across ___ quarters of Skyloom's own history.
Show hint
Look at "Let's learn."
Show answer
340 deals, across 6 quarters. A real, sizeable sample, and still too sparse for the model to reliably beat a simpler, hand-tuned method.
Short answer, where it wouldn't matter
4. Describe a situation where this same memo structure would NOT need to recommend pausing an AI initiative.
Show hint
Think about what the backtest result would need to show instead.
Show answer
Model answer: If the backtest had shown QuotaSense beating the existing method's accuracy, the same LEAD structure would support recommending the rollout proceed, the link step is what decides the direction, not a default stance against AI initiatives.
Short answer, apply it yourself
5. Think of a time you had to give feedback that contradicted a decision someone senior to you had already publicly committed to. What real evidence, if any, backed your recommendation?
Show hint
Think about whether your feedback could be checked against something concrete, or whether it was mostly a judgment call.
Show answer
Model answer: A team lead once championed switching vendors based on a sales demo; pulling actual usage logs from a two-week trial showed the new vendor's tool was slower for the team's real workflows, a concrete number that made the pushback land, not just a preference.
Short answer, work the number
6. If QuotaSense's average error had been 30 percent instead of 34, just barely beating the pipeline's 31 percent, would the memo's recommendation likely change?
Show hint
Think about how close a margin needs to be before other factors, like cost, start to matter more.
Show answer
Model answer: Possibly, but the memo would likely still flag the real cost gap, $180,000 plus $24,000 a year against a nearly free existing method, as worth weighing against such a small accuracy edge, rather than an automatic green light.
Before you close the answer
Why this works
Tests whether you can deliver hard, evidence-based pushback to someone senior without either softening it into uselessness or turning it into a personal conflict, and whether the evidence itself is genuinely AI-specific rather than a generic budget objection.
Follow-up traps
"What if the CEO just overrules the memo anyway?" Response: the backtest data and the concrete two-quarter plan still exist either way, ready to be pointed to the moment the rollout underperforms, which protects the team and the decision-making record regardless.
"Isn't three points of forecast error pretty small?" Response: three points on a company-wide sales forecast is a real, material gap, and it's paired with a real $180,000-plus cost, on a method that already works for less.
If pressed
The backtest deliberately included Q5, a seasonally unusual quarter, specifically because a model trained mostly on typical quarters often degrades hardest exactly when the pattern shifts, which is exactly what QuotaSense's 41 percent spike in that quarter confirmed.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.