ConceptIntermediateAI Opportunity & Model Strategy / Feasibility assessment and technical spikes / #22
When should you skip the spike and just build?
ORDER·the quarter engineers started building their own smoke tests
Kestrel Waste Systems runs residential collection routes for about forty towns. BinSense is its feature that predicts which bins will overflow before the truck arrives, so routes can adjust the same morning. Priya Wexler is the AI PM who has to decide, every quarter, which features actually need a full spike before they ship.
The direct answer
Skip the formal spike when the decision is cheap to reverse, touches a small blast radius, and you already have a proven model or a fast canary that would catch a bad call within hours. Run the full spike when a wrong call is expensive to undo, affects many routes or people at once, or is genuinely new territory with no track record behind it. The line is reversibility and blast radius, not how much the feature sounds like "real AI."
Do this, in order
Ask first whether the decision is cheap to reverse.Why: a cheap-to-reverse call is safer to canary and watch than to spend weeks proving in advance.
Ask how many routes or people one wrong call touches.Why: a small blast radius needs far less upfront proof than a fleet-wide one.
Check whether you already have a proven model or pattern from a past feature.Why: re-spiking something you've already validated once wastes the exact time a spike is meant to save.
Write down, in one line, what a fast canary would catch and how soon.Why: if a canary catches a bad call in hours, a six-week spike is buying you very little.
Reserve the full spike for decisions that are genuinely new, hard to reverse, or wide in blast radius.Why: that's where a wrong guess actually costs something a canary can't catch in time.
Say plainly when the safer choice really is a full spike, like reassigning a whole day's routes with no rollback.Why: shows judgment about where the rigor earns its cost, not a blanket rule applied everywhere equally.
How to answer this, stage by stage
Nobody is scoring whether you know that spikes are good practice. They're scoring whether you can say, out loud, exactly when one is a waste of six weeks.
Stage 1
Scope it to one real decision
Say it like this
"Let's ground this in BinSense at Kestrel Waste Systems, and the actual quarter where the team's own spike rule started costing more than it protected."
Why this works
Keeps the answer from becoming a general essay about "moving fast versus being careful."
Stage 2
Say the structure out loud
Say it like this
"I'll run this as ORDER. Outcome, what we're protecting. Reversibility, what's hardest to undo. Dependency, what's forced anyway. Evidence, what's cheap to learn first. Rank, the actual call."
Why this works
Signals a repeatable way to decide, not a gut feeling about which features "feel risky."
Stage 3
Reframe: it isn't "does this touch a model," it's "how hard is it to take back"
Say it like this
"This isn't really a question of whether a feature involves AI. It's a question of what happens if the first version is wrong, and whether that's a five-minute fix or a fleet-wide mess."
Why this works
This is where a strong answer separates from someone who says "always spike anything with a model in it."
Stage 4
Give the rank
Say it like this
"Flagging a bin that's reported full three times in a row for a manual reroute: skip the spike, ship behind a canary, watch it for a week. Automatically reassigning an entire day's routes based on a fill-level model: run the full spike, because a bad guess there strands real trucks on real streets."
Why this works
This is the direct answer, applied to two concrete candidates instead of stated as an abstract rule.
Stage 5
Prove it with the compressed failure
Say it like this
"We didn't have that rule written down. So four small, easily reversible tweaks last year each waited the same six-week spike as the fleet-wide feature. About nineteen weeks were lost to spikes that never changed a single decision, and engineers started quietly building their own before-and-after smoke tests just to move faster than the process allowed."
Why this works
Compresses the whole case into the one number a written reversibility rule would have prevented from ever accumulating.
Stage 6
Name the AI-specific reasoning, then close
Say it like this
"The honest reason this isn't just generic process advice is that a model's mistakes are probabilistic, not deterministic, so the real question is never 'could it be wrong,' it's always 'how fast can we see it and undo it.' I wouldn't skip the spike for the fleet-wide auto-reroute feature, since a bad day there strands trucks with no fast rollback. For a reversible tweak, the canary is the spike."
Why this works
Closes with real judgment about where the concern still applies, and restates the direct answer in one breath.
Let's learn
Here is what happens when a process built for one hard decision gets applied to every decision, hard or not.
Before BinSense, a driver reported a bin as full over the radio, and a dispatcher manually slotted an extra stop into the day's route, about ten minutes of dispatcher time each time it happened, a few times a week. With BinSense's core model, fill-level predictions come back automatically, and the harder question becomes what the system is allowed to do with them on its own.
The five letters, held up as one page. Reversibility is the step this question is really testing.
Here's the turn: the spike process at Kestrel was built after one bad launch that shipped with no validation at all. The fix that followed wasn't a rule about which decisions needed proof first. It was a blanket six-week spike for every feature touching a model, regardless of how easy that feature would be to undo if it went wrong.
What a standard spike actually costs, broken down
The same 5.5 weeks ran for a fleet-wide routing change and for a bin-flagging tweak a canary could have proven in three days.
At its worst, a reversible, low-stakes tweak sits in a six-week queue behind decisions that actually deserve that scrutiny, engineers get tired of waiting, and they start quietly building their own before-and-after checks outside the process entirely, with nobody reviewing those checks at all.
The process wasn't protecting Kestrel from bad decisions anymore. It was just protecting the calendar from ever being asked which decisions actually needed protecting.
The choice I would take back
Nobody ever wrote down a rule for what does and doesn't need a full spike. That made sense right after the bad launch, when the instinct was simply "never skip validation again." It stopped making sense the moment "never skip it" quietly became "never think about whether this specific decision needs it."
What I would leave alone: I wouldn't relax the full spike for the fleet-wide auto-reroute feature itself, since a bad version of that one strands real trucks on real streets with no fast way back.
The lesson: a spike is a tool for a specific kind of decision, the hard-to-undo, wide-blast-radius kind. Applied to every decision equally, it stops being caution and starts being a tax on the easy ones.
Now here is the same thing as a story
The short version above is what you'd say defending a lighter process in a planning review. Read this one for how the rule quietly hardened over four quarters with nobody deciding it on purpose.
Every dispatch shift at Kestrel starts the same way: a printed route sheet, a radio channel, and a coffee that's always gone cold by the second call.
The second step is where the same five weeks ran for every feature, no matter how easy it would be to undo.
Two years ago, a route-scoring feature shipped with no spike at all, misjudged traffic patterns badly for a week, and cost the depot real overtime before anyone rolled it back. After that, the rule became: every feature touching a model gets the full process, no exceptions.
Nobody decided, on any single day, to spike everything equally. It just never stopped being the default after the one time it was needed most.
By the third quarter, a genuinely small, cheap tweak, flagging any bin reported full by a driver three times running for a manual reroute, went into the same six-week queue as the fleet-wide auto-reroute model. It came out the other side unchanged, because there had never really been much to learn from spiking it in the first place.
Knowledge spark: why would a small AI feature need less proof than a big one?
A model's mistakes are a matter of probability, not certainty, so the real safeguard is how fast you can see a bad call and undo it, not how much upfront proof you gather. A cheap-to-reverse, narrow-impact decision can lean on a fast canary and quick rollback. A hard-to-reverse, wide-impact one can't, because by the time you'd notice, the damage is already spread across the whole fleet.
Two doors. One swings back easily. One is bolted shut the moment the trucks are already moving.
By the fourth quarter, a senior engineer named Tobin Marsh had quietly built his own before-and-after comparison script that he ran on every small model tweak before it ever entered the formal process, just so he could tell, in an afternoon, whether six weeks of waiting was even worth proposing. Nobody reviewed his script. Nobody knew it existed except the engineers using it.
The two trivial features sat in the same corner. The one that actually needed the full spike sat alone, far from both of them.
The real question was never whether BinSense's model deserved scrutiny. It was whether every single decision touching that model deserved the same six weeks of it.
Four questions that were never written down anywhere the team could point to.
When the blanket rule was first adopted, someone said, "let's just require the full process for anything touching the model, so we never get burned like that again," and it sounded reasonable, since the one bad launch was still fresh in everyone's memory.
Cumulative weeks lost to unnecessary full spikes, by quarter
Nineteen weeks, over one year, spent proving things a canary would have shown in days.
Rerun the same year with a written reversibility rule in place: the bin-flagging tweak and the threshold adjustment both ship behind a three-day canary instead of a six-week spike. The fleet-wide auto-reroute feature still gets the full process, since a bad version of it really does strand trucks with no fast way back. Nineteen weeks return to the roadmap, and Tobin's private script becomes an official, reviewed part of the canary step instead of a quiet workaround.
What I'd tell myself, hearing that an engineer had built his own validation process on the side: the rule was never wrong to exist. It was wrong to apply to everything the same way, once "never again" hardened into "always, regardless."
ORDER, the rule that finally noticed the differenceNot a script for skipping validation whenever it's inconvenient. ORDER is what tells you exactly which decisions actually earn the six weeks.
O
Outcome. What all candidates compete to move.
Keeping routes efficient without ever stranding a truck or a bin on a wrong model call.
Without a shared outcome, ranking which features need a spike is just opinion.
R
Reversibility. Which decision is hardest to undo.
A bin-flagging tweak reverses in minutes. A fleet-wide auto-reroute, once trucks are dispatched on it, does not.
This is the hardest step, and the one the blanket rule never actually asked.
D
Dependency. What unblocks what.
The auto-reroute feature needed the core fill-level model validated first. The bin-flagging tweak depended on nothing else in flight.
Some order is forced by what the roadmap actually requires, not by preference.
E
Evidence. What's cheap to learn first.
A three-day canary on the bin-flagging tweak would have told the team everything the six-week spike eventually confirmed anyway.
Cheap evidence, gathered first, is what makes a full spike optional for the right candidates.
R
Rank. State the order, defend the top.
Full spike first for the auto-reroute feature, canary-only for the bin-flagging tweak and the threshold adjustment, in that order of urgency.
The ranking follows directly from reversibility and blast radius, not from which feature sounds more like "real AI."
The recap, one line per letter: outcome is routes staying efficient without a stranded truck, reversibility is a bin flag reversing in minutes against a fleet reroute that doesn't, dependency is the core model needing validation before the reroute feature can, evidence is a three-day canary proving what a six-week spike eventually confirmed anyway, and rank is spiking the hard-to-undo feature first and canarying the rest.
And if you want to be sure it really works, try it somewhere elseSame five letters, an HVAC field-service company instead of a waste fleet. Different flip family entirely, the same missing rule.
Doreen Kessling runs product at Cinder & Vale HVAC Services, where one feature looks up a diagnostic code from a unit's error display, and a newer one predicts which units are likely to fail within the next two weeks so a technician gets dispatched early. Mapped onto ORDER: outcome is fewer emergency no-heat calls without wasting technician dispatches. Reversibility says the diagnostic lookup is trivially reversible, wrong once, corrected next time, no truck involved. The predictive-failure alert is not: once a truck rolls on a false alarm, that technician's whole morning is gone. Dependency says the lookup needs nothing else validated first. Evidence favors a small canary for the lookup and a real spike for the predictive alert. But the flip here is an abandonment one, not a workaround: after a few early false alarms wasted real truck rolls, technicians didn't build a private process, they just quietly stopped opening the predictive alert feed at all, with no complaint filed anywhere.
A different flip entirely: not a private workaround built in the shadows, but a feed that quietly stopped getting opened at all.
Swap the trigger and it still runs.
Speed: an interviewer caps you at sixty seconds. Say "skip the spike when it's cheap to reverse and small in blast radius, spike it when it isn't, that's the whole rule," and stop.
Cost: no time to build a formal reversibility rubric before the next feature ships. Say so honestly, and ask the one question, "how fast could we undo a wrong call here," before defaulting to either extreme.
The feature turns out to be lower stakes than assumed, for real: if a "risky-sounding" feature actually has a fast, cheap rollback built in, that's a legitimate reason to skip the full spike, not a shortcut being taken.
Where people run it wrong.
They treat "touches a model" as the trigger for a full spike, instead of reversibility and blast radius.
They apply one blanket process to every decision after getting burned once, instead of writing down what actually needs the extra scrutiny.
They let engineers build private, unreviewed shortcuts because the official process has stopped fitting most of the actual decisions.
How to use it live. The moment an interviewer asks when to skip a spike, ask yourself: how fast and how cheaply could we undo this if it's wrong, and how many people or routes would feel it in the meantime? Answer those two, and the rest of the call makes itself.
Flashcards (tap any card to flip it)
1 · THE FLIP FAMILY
What flip family is this?
Tap to flip
ANSWER
Workaround flip: with no shared rule for what needed a spike, an engineer built his own private before-and-after script and ran it outside the official process entirely.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Priya Wexler, the AI PM at Kestrel Waste Systems, who had to decide which BinSense-related features actually needed a full spike.
3 · THE HABIT
What did the team stop doing after the one bad launch two years earlier?
Tap to flip
ANSWER
They stopped asking whether a given feature actually needed a full spike, and started requiring the same six-week process for every feature touching a model, without exception.
4 · THE FLIP, IN THIS STORY
What's the two setting switch here?
Tap to flip
ANSWER
Waiting out the official six-week process versus quietly building a private, unreviewed shortcut. No middle setting once the process stopped fitting most real decisions.
5 · THE OLD DECISION
What decision would you take back?
Tap to flip
ANSWER
Never writing down a rule for what does and doesn't need a full spike, leaving "never skip validation again" to harden into "always, regardless of the decision."
6 · THE NUMBER
Fill in the blank: over one year, unnecessary full spikes on cheap-to-reverse features cost about ___ weeks.
Tap to flip
ANSWER
Nineteen weeks.
7 · THE REPLAY
Same year, a written reversibility rule in place from the start. What changes?
Tap to flip
ANSWER
The bin-flagging tweak and threshold adjustment ship behind a three-day canary. The auto-reroute feature still gets the full spike. Nineteen weeks return to the roadmap, and the private script becomes an official part of the canary step.
8 · CROSS PRODUCT TRANSFER
Section 4 answers this same question again for a different product, with a different flip family. Which product, and which family?
Tap to flip
ANSWER
Cinder & Vale HVAC Services' predictive-failure alert. The flip is abandonment: technicians quietly stopped opening the alert feed after a few false alarms, with no complaint ever filed.
Check yourself Score: 0 / 0
True or false
1. True or false: this answer argues that every feature involving a model should skip formal validation to move faster.
True
False
Show hint
Look at "what I would leave alone" and the auto-reroute feature.
Show answer
False. The fleet-wide auto-reroute feature still gets the full spike. Only decisions that are cheap to reverse and narrow in blast radius get to skip it.
Multiple choice
2. What actually determines whether a feature needs a full spike, according to this answer?
A. Whether the feature uses a large language model or a simpler classifier.
B. How hard a wrong call is to reverse, and how many people or routes it affects at once.
C. How excited the team is about the feature.
D. How much engineering time is available that particular quarter.
Show hint
Look at the direct answer and the "reversibility" step.
Show answer
B. Reversibility and blast radius are the two things that actually decide whether the extra weeks of a full spike are worth their cost.
Fill in the blank
3. Fill in the blank: Kestrel's standard spike process took about ___ weeks, broken into data pull, curated set, eval build, and review meetings.
Show hint
Look at the stacked bar chart, "what a standard spike actually costs."
Show answer
5.5 weeks. The same 5.5 weeks ran for the fleet-wide feature and for a tweak a three-day canary could have proven instead.
Short answer, where it wouldn't matter
4. Name a feature where the full six-week spike genuinely was worth keeping, and say why.
Show hint
Look at "what I would leave alone."
Show answer
Model answer: The fleet-wide auto-reroute feature. A bad version of it strands real trucks on real streets with no fast rollback, which is exactly the hard-to-undo, wide-blast-radius case a spike exists for.
Short answer, apply it yourself
5. Think of a small change your team or workplace treats with far more process than it needs. What would a written rule about reversibility have let you skip?
Show hint
Think about a decision that's easy to undo but still gets the same heavy review as a hard-to-undo one.
Show answer
Model answer: A minor email-template wording change requiring the same multi-person approval as a pricing change, when a quick send-and-watch-the-open-rate test would have answered the question in a day.
Short answer, work the number
6. If Kestrel adds a written reversibility rule and cuts two of the four unnecessary spikes next year instead of all four, roughly how many weeks would that save?
Show hint
Look at how the 19 weeks broke down across roughly four instances over the year.
Show answer
Model answer: Roughly 9 to 10 weeks, about half of the 19 weeks lost, since each unnecessary spike cost close to 5 weeks over what a fast canary would have needed.
Before you close the answer
Why this works
Tests whether you'll apply judgment about reversibility and blast radius to each decision, or default to either "always spike" or "never spike" out of habit.
Follow-up traps
"Isn't 'cheap to reverse' just an excuse to skip rigor whenever it's convenient?" Response: no, because the fleet-wide feature, the genuinely hard-to-undo one, still gets the full spike under this rule. The line moves with the actual risk, not with convenience.
"What if a 'reversible' feature turns out to have a hidden, expensive failure mode?" Response: that's exactly what the fast canary is for, catching it in hours instead of assuming reversibility was correctly judged and never checking again.
If pressed
The written rule that followed this incident sets a specific bar: any feature where a wrong call could still be corrected within 24 hours for under 5 percent of affected routes qualifies for canary-only validation, everything else defaults to the full spike.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.