ConceptFoundationalAI Opportunity & Model Strategy / When NOT to use AI / #2
Explain why a deterministic rules engine sometimes beats a model, with an example.
PICK · Stagbrook's own formula held one number steady for fourteen months. The model gave that same steady item four different answers in ten weeks
Stagbrook Distribution supplies hardware and fasteners to about 140 independent stores. Tanvir Alkemade owns the reorder logic for its 4,200 SKUs. Baastian Rourke, a forecasting engineer, built DemandSync, a model that recalculates each item's reorder point every week. When leadership rolled it out to every SKU at once, a box of 3/8-inch lag screws, an item that had barely moved in over a year, started getting a different reorder number almost every week.
The direct answer
Use a rules engine, not a model, whenever the actual decision logic is fully known and stable enough to write down as explicit conditions today. At Stagbrook, the reorder point for a steady item like the lag screws is just lead time times average daily demand, plus a safety-stock cushion, math that hasn't changed in years. Routing it through a weekly-retrained model didn't make the number smarter. It made a known answer swing by nearly three times, for no reason at all.
Do this, in order
Route any decision with fully known, stable logic straight to a rules engine, not a model.Why: this is the whole call, and skipping it is exactly what turned a solved problem into an expensive one.
Before building anything, ask one question: can today's logic be written down completely and correctly as explicit rules?Why: this is the kill criteria, the single test that decides everything after it.
Save model complexity for demand that's genuinely fuzzy and pattern-based, not for logic you already have exact math for.Why: names where the model actually earns its keep, so the answer isn't "never use one."
Treat an unexplained swing in a "known" number as the signal that a decision got wrongly routed to a model.Why: this is the guardrail that catches the failure in production, before it costs real money.
Weigh a model's ongoing retraining and monitoring cost against a formula's near-zero, one-time cost, before choosing.Why: this is the cost asymmetry, the actual reason the mistake is cheap in one direction and expensive in the other.
Don't just smooth a misapplied model's output to make the swings look smaller. Remove it from logic it was never needed for.Why: smoothing hides the mismatch instead of fixing it, and it was the alternative Stagbrook actually rejected.
How to answer this, stage by stage
Nobody is grading whether you sound impressed by models or suspicious of them. They're grading whether you can tell a fully known decision from a genuinely fuzzy one, on the spot.
1
Pin it to one number-driven decision, not the whole product
Say it like this
"Let's ground this in one real SKU. Stagbrook Distribution is a wholesale hardware supplier, and Tanvir owns the reorder point for about 4,200 items, including a box of 3/8-inch lag screws that sells about 42 boxes a day, steady as anything."
Why this works
Keeps the interviewer grading a real decision, not a theory about rules versus models in general.
2
Name the method up front
Say it like this
"I'll use PICK. Position, when a rules engine actually beats a model. Impact, what's lost using a model where rules would do. Cost asymmetry, which mistake is cheap and which is expensive. Kill criteria, the one test that decides it."
Why this works
Two seconds of structure tells the room a method is coming, not a mood.
3
State the position plainly, before any story
Say it like this
"My position: when the actual decision logic is fully known, stable, and can be written down today as explicit conditions, a rules engine beats a model. It's consistent, instantly explainable, and cheap to check. The model would just be learning an approximation of math you already have exactly."
Why this works
This is the direct answer, said early enough that nothing after it can blur it.
4
Tie the line to something only a model raises
Say it like this
"This only matters because the alternative is a model that retrains every week. A rules engine runs the same formula on the same inputs every time. A model can drift, retrain on noise, and hand you a different answer for a SKU where nothing in the real world actually moved."
Why this works
Keeps the answer anchored to model behavior, not generic build-it-simple advice.
5
Bring the real numbers
Say it like this
"Here's what happened. The formula had held the lag screws' reorder point at 272 boxes for fourteen months, moving only once, when the supplier's own contracted lead time changed. Once DemandSync took over every SKU, that same reorder point swung between 180 and 460 across ten weeks, with no change in demand and no change in lead time."
Why this works
A real number the whole argument would fall apart without, not a hypothetical.
6
State the trade, then hand over the test
Say it like this
"Keeping the formula for known logic costs almost nothing, maybe a few hours a year. Routing it through the model instead cost about six hours of retraining time a week, plus a rush shipment when the number ran too low and warehouse space tied up when it ran too high. My test: can today's logic be written down completely as rules? If yes, that's a rules-engine decision. Save the model for the ones you genuinely can't write down, like a snow shovel's reorder point ahead of a forecast cold snap."
Why this works
Ends on something countable and hands over a test instead of a feeling.
Let's learn
For fourteen months, one number didn't move: 272 boxes of 3/8-inch lag screws, the exact point where Stagbrook reordered more.
Stagbrook Distribution is a wholesale hardware supplier, and a formula decides when each of its 4,200 SKUs gets reordered before the shelf runs dry. The formula is old and boring, in the best way: average daily demand, times lead time, plus a safety-stock cushion for the normal ups and downs. For the lag screws, that math held reorder point at 272 boxes for fourteen straight months, moving only the one time the supplier's own contracted lead time changed.
One side gives the same answer every time it's asked. The other side is trained, and a trained thing can drift.
Then DemandSync, a forecasting model retrained every week on sales, weather, and local promotions, took over reordering for every SKU at once, including the screws. Ten weeks later, that same reorder point had swung between 180 boxes and 460 boxes, with nothing in the real world moving at all.
Here's the important part. The swings themselves weren't really the problem. The problem is what Tanvir's floor team did next: they stopped trusting the number and started phoning the supplier to double-check every reorder by hand, on every steady item, because they couldn't tell a real signal from noise anymore.
The lag screws never needed a model. They needed someone to leave the math alone.
Knowledge spark: what is safety stock?
A buffer of extra units added on top of plain lead-time demand. It's sized to how much daily demand actually swings, so a normal bad week doesn't turn into a stockout. For a steady item, that buffer barely moves, because the swings it's covering for barely move.
What it cost at its worst: one week the model's number ran low, and stock nearly ran out, forcing a $340 rush shipment. Another week it ran high, and left almost $2,150 in screws sitting on a pallet for eleven weeks, crowding out faster-moving spring stock. The model didn't make ordering smarter. It made a fully solved problem expensive again.
Cost, by the numbers: keeping the formula versus keeping the model, per quarter, for the 900 stable SKUs
Cheap, foreverExpensive, and it bought nothing
The rules-engine number covers an occasional review of the formula's own inputs. The model number covers about six hours a week of retraining and monitoring time, plus the rush shipment and the pallet space tied up by the swings themselves.
The choice I would take back
Every SKU got routed through DemandSync so the company would have "one system." I would take that back for the roughly 900 SKUs whose demand and lead time barely move, the lag screws among them. Their reorder logic was already fully known. The model wasn't learning anything there. It was adding noise to math that was already solved.
What I would leave alone: snow shovels are a different story. Their real reorder point genuinely needs to jump from 20 units in July to 300 units in November, ahead of a forecast cold snap, and no fixed formula can see that coming. DemandSync earns its cost there. I'd leave it running exactly as it is for every SKU whose demand actually depends on something the model can see and a formula can't.
The lesson: a model is worth its retraining cost only where the logic is genuinely fuzzy. When you already have the exact math, written down in years of practice, building a model to guess at it again doesn't make the number smarter. It just makes a known answer start moving for no reason, and now someone has to figure out whether to trust it.
Now here is the same thing as a story
The short version above is what you'd actually say out loud. Read this one for what it cost Stagbrook to learn it the slow way.
Every Monday morning, Tanvir Alkemade ran the same six-line report before anyone else was in the building. Reorder points for 4,200 SKUs, checked against actual stock on hand, no surprises. He'd owned that formula for six years, and for most of that time it had been the least interesting part of his job. That was the whole point of it.
DemandSync arrived first for the fun items: snow shovels, patio heaters, the string trimmers that sell out in a week every April. Baastian Rourke had built it to catch exactly what a fixed formula couldn't, weather forecasts, local promotions, the way one warm February changes a whole spring order. For those SKUs, it worked. Snow shovel stockouts, a real problem the winter before, mostly disappeared.
Twelve months of the model doing real work, and then one decision widened its job past what it was built for.
So when leadership asked why Stagbrook still ran two different systems for reordering, one for the interesting items and the old formula for everything else, the answer felt obvious: give DemandSync all 4,200 SKUs. One system. Cleaner to explain, cleaner to maintain, one dashboard instead of two.
Nobody flagged the lag screws specifically. Why would they. They were the most boring item in the warehouse.
Five weeks into the full rollout, the screws' reorder point posted at 460 boxes, almost double the number the formula had run for over a year. The warehouse ordered accordingly. Three weeks after that, it posted at 180, and stock dropped to 34 boxes with a delivery still six days out, well below what the site actually needed to stay covered. A rush shipment closed the gap for $340 in freight nobody had budgeted.
Five steps, and the whole argument sits inside the middle one: a number that used to sit still, jumping for no reason anyone on the floor could name.
The floor supervisor was the one who finally said it out loud, standing next to a pallet of screws nobody had asked for: "Why do we keep almost running out of these things? Nothing's changed out here. Same customers, same trucks, same everything."
We didn't lose to bad math. We lost to math that was never supposed to move, moving anyway.
Tanvir pulled ten weeks of DemandSync's predictions for the screws and lined them up next to the old formula's answer. Week to week: 260, 410, 190, 305, 460, 220, 275, 180, 390, 300. The formula's own number, recalculated fresh from the exact same sales data, sat at 272 every single time, because nothing about demand or lead time had actually changed all quarter.
Before pulling the screws off DemandSync entirely, the team considered a smaller fix: cap how much the model's output could move week to week, or average its last four predictions instead of trusting the newest one. It looked like a way to keep "one system" and calm the swings at the same time. They rejected it. Damping the model's output doesn't make it correct, it just makes a wrong number harder to notice, and Stagbrook would still be paying to retrain a model on logic that never needed learning in the first place. A smoothed 320 is still not the formula's checkable 272.
Here's the replay. Same rollout, same push for "one system." This time, before flipping every SKU to DemandSync, someone runs one extra check first: for each SKU, does its actual sales history look steady enough that the formula's own output barely moves month to month? For the roughly 900 SKUs where the answer is yes, including the lag screws, they stay on the formula. DemandSync keeps the other 3,300, the genuinely seasonal, promotion-driven, weather-sensitive items it was actually built for. That one afternoon of sorting SKUs by how steady their demand really is costs about $400 a quarter, forever, to keep current. It also means the floor supervisor never has to ask why a box of screws keeps almost running out.
What I'd tell myself, sitting in the meeting where "one system" first sounded like the obvious call: consistency across every SKU is not the same goal as correctness for any one of them. Sometimes the cleanest-looking decision is the one that quietly stops asking whether each SKU actually needed the tool it just got handed.
PICK, and the test that keeps a model off math you've already solved
Not a case against models. DemandSync earned its cost on the seasonal items. PICK is what separates that real win from the mistake sitting one aisle over.
One side is a formula anyone can check by hand in thirty seconds. The other side costs real money to relearn a question that was never actually open.
PPosition. When a rules engine actually beats a model.
When the real decision logic is fully known, stable, and can be written down today as explicit conditions, a rules engine gives perfectly consistent, instantly explainable, cheap-to-check behavior that a probabilistic model can't match. The model would just be learning an approximation of logic you already know exactly.
For the lag screws, that logic is: average daily demand times lead time, plus a safety-stock cushion for normal swings. Nothing about that math was ever uncertain.
State the position before any story, so it doesn't read as reverse-engineered from what already went wrong.
IImpact. What's lost using a model where rules would do.
Unnecessary unpredictability on a decision that was never actually uncertain. Harder-to-explain output, since nobody on the floor can say why 272 became 460 when nothing changed. And an ongoing retraining and monitoring cost for logic that never needed to be learned in the first place. At Stagbrook that meant a floor team that stopped trusting a number it used to take for granted, on top of a real $340 rush fee and $2,150 in tied-up stock.
CCost asymmetry. The heart of it.
Building and keeping a rules engine for logic that's actually fully known and stable is cheap: write the formula once, review it occasionally, done. Building and keeping a model for that same logic adds real ongoing cost, retraining time, monitoring, infrastructure, and buys real unpredictability in return, not accuracy. Rules for the 900 stable SKUs ran about $400 a quarter. The model for the same 900 SKUs ran about $17,300 a quarter, once retraining hours, the rush shipment, and the tied-up pallet space were all counted. Start from the cheap mistake. Only pay for model complexity where the logic genuinely isn't known yet.
KKill criteria. The one test.
Can the actual decision logic be written down completely and correctly as explicit if/then rules, today? If yes, a rules engine is very likely the better choice. Model complexity should be reserved for genuinely fuzzy, pattern-based judgment that can't be fully specified as rules, like a snow shovel's reorder point ahead of weather nobody can write into a formula. For the lag screws, the answer was yes the whole time. Nobody had asked the question.
Two questions, not one gut call: how steady is the demand, and how well does a fixed formula already track it.
The kill line, charted: DemandSync's weekly reorder point for the lag screws against the formula's own answer
The formula's answer, unchangedThe model's answer, ten different guesses
The formula's line is flat because nothing about the screws' demand or lead time ever moved. Every point where the model's line strays from it is pure retraining noise, not a real signal anyone on the floor could act on.
The trade worth saying out loud: a rules engine can't adapt to a genuinely new pattern its formula doesn't capture, and that's a real cost when the logic is truly fuzzy, which is exactly why DemandSync still runs the seasonal SKUs. A model can adapt to real, changing patterns, but it costs ongoing retraining and monitoring, and that cost is actively harmful when it's spent on a SKU whose real answer never moves. Pay the model's cost only where the pattern is genuinely worth learning. Everywhere else, that same cost buys nothing but noise.
And if you want to be sure it really works, try it somewhere else
Same four letters, a meal-kit delivery app's refund desk instead of a hardware warehouse, and the unproven belief is that "AI-powered" beats a stopwatch.
TrailFeast is a meal-kit delivery service, and its refund policy for late orders is one clean rule: if a delivery lands more than 10 minutes past the promised window, the customer gets an automatic refund. Someone on the team built a model instead, weighing weather, driver history, and order size to "predict" whether a late order deserves a refund.
Position: whether an order was late is fully known logic, a timestamp comparison against a promise TrailFeast already made. A rules engine beats a model here every time. Impact: the model doesn't just add unpredictability in the abstract, it actively denies refunds that were plainly owed, because it's weighing signals that have nothing to do with the actual promise made to the customer. Cost asymmetry: the rule costs almost nothing to run, a subtraction and a comparison. In a sample of 300 clearly-late orders, the model wrongly denied 66 of them, 22 percent, each needing about six minutes of a support agent's time to override, close to $260 a week in support hours the rule would have spent for free. Kill criteria: can "was this order late" be written down completely as a rule today? Yes, so it goes to the rules engine. The model gets reassigned to something it's actually suited for: spotting a pattern of many refund claims from one account over time, a judgment call no single timestamp can make.
Three of the four branches never needed a model at all. The fourth one is where it actually earns its keep.
Swap the trigger and it still runs.
Speed: an interviewer caps you at ninety seconds. Skip straight to the kill test: can today's logic be written down completely as rules, yes or no.
Cost: no time to check whether the logic is really fixed before a decision has to ship. Fine, but say the test out loud and make whoever owns the roadmap answer it, instead of letting the deadline default to "add a model, it looks safer."
The model got better, for real: say a later version of DemandSync only updates when a SKU's actual sales pattern moves outside a known band, instead of retraining every week regardless. That's the moment it earns a second look on stable SKUs too, because now it behaves like a rules engine with a safety net, not a source of its own noise.
Where people run it wrong.
They assume "model" always means smarter, and never stop to ask whether the logic was ever actually unknown.
They treat one bad week from a model as an accuracy problem to retrain away, instead of an unpredictability problem to route away entirely.
They keep a model on a decision they've already proven is fixed, because turning it off feels like a step backward.
How to use it live. If you're ever asked whether a decision needs a model at all, buy yourself a second with one plain question, said out loud: "can I write the actual rule down right now?" If you can, say it. That question is the whole method, asked instead of stated.
Flashcards (tap any card to flip it)
1 · THE FRAMEWORK
What framework fits a question asking you to choose between two ways of solving the same decision?
Tap to flip
ANSWER
PICK: state the position, name the impact on each side, find which mistake is cheap versus expensive, then give the one test that actually decides it.
2 · THE PERSON
Who is this answer about?
Tap to flip
ANSWER
Tanvir Alkemade, who owns reorder-point logic for about 4,200 SKUs at Stagbrook Distribution, a wholesale hardware and fastener supplier.
3 · THE POSITION
What's the real position on rules engines versus models here?
Tap to flip
ANSWER
When decision logic is fully known, stable, and can be written down today as explicit rules, a rules engine beats a model. It's consistent, instantly explainable, and cheap to check. A model there would just be learning an approximation of math you already have exactly.
4 · THE GAP
What's the two-number gap this whole answer turns on?
Tap to flip
ANSWER
The formula held the lag screws' reorder point at 272 boxes for fourteen months. Once the model took over, that same number swung between 180 and 460 across ten weeks, with nothing in the real world actually changing.
5 · THE REVERSAL
What old decision would Tanvir take back?
Tap to flip
ANSWER
Routing every SKU through DemandSync so the company would have "one system." That made sense as a simplicity goal, but it was the wrong call for the roughly 900 SKUs whose demand and lead time barely move.
6 · THE NUMBER
Fill in the blank: keeping the rules engine for the 900 stable SKUs cost about $___ a quarter. Routing them through the model instead cost about $___ a quarter.
Tap to flip
ANSWER
About $400. About $17,300, once retraining hours, the rush shipment, and the tied-up pallet space are all counted.
7 · THE KILL TEST
What's the one test for choosing a rules engine over a model?
Tap to flip
ANSWER
Can today's decision logic be written down completely and correctly as explicit if/then rules? If yes, a rules engine is very likely the better choice. Save model complexity for logic that genuinely can't be fully specified as rules.
8 · CROSS-PRODUCT TRANSFER
Section 4 answers this same question for a different product. Which one, and what plays the role of the lag screws' formula?
Tap to flip
ANSWER
TrailFeast, a meal-kit delivery service. The role goes to its late-delivery refund rule, a plain timestamp comparison, logic just as fully known as the lag screws' reorder math.
Check yourself Score: 0 / 0
True or false
1. True or false: the real problem with DemandSync on the lag screws was that its predictions were less accurate than the formula's.
True
False
Show hint
Ask what actually changed for the floor team: a bad prediction, or an unpredictable one.
Show answer
False. Accuracy on average wasn't the issue. The issue was that a known, checkable answer started moving for no reason, which is a stability and trust problem, not a plain accuracy problem.
Fill in the blank
2. Fill in the blank: the formula held the lag screws' reorder point at ___ boxes for fourteen months. Once DemandSync took over, that same number swung between ___ and ___ across ten weeks.
Show hint
Look at stage 5 of the walkthrough, and the flashcard called "the gap."
Show answer
272. 180. 460. That gap, a known, checkable number swinging by nearly three times, is the whole argument.
Multiple choice
3. Which SKU is the better candidate for keeping DemandSync, instead of switching to the rules engine?
A. 3/8-inch lag screws, because they're the highest-volume item in the warehouse.
B. Snow shovels, because their real reorder point genuinely depends on weather a fixed formula can't see.
C. Whichever SKU the warehouse floor complains about least.
D. Whichever SKU has the highest dollar value per box.
Show hint
Ask which item's real-world demand can't be written down as a fixed rule.
Show answer
B. Snow shovels need a reorder point that genuinely moves with a forecast, from 20 units to 300, something no fixed formula can capture. That's exactly the fuzzy, pattern-based judgment a model is for.
Short answer, name the rejected alternative
4. What alternative did Tanvir's team consider instead of pulling the stable SKUs off DemandSync entirely, and why was it rejected?
Show hint
Look near the replay in the story section, right before "here's the replay."
Show answer
Model answer: Capping how much the model's weekly output could change, or averaging its last few predictions, to calm the swings while keeping "one system." Rejected because damping a wrong number doesn't make it correct, it just hides the mismatch, and Stagbrook would still be paying to retrain a model on logic that never needed learning.
Multiple choice
5. Which statement best matches the position on rules engines versus models for known, stable logic?
A. Always prefer a model, since it can only get smarter with more data over time.
B. Always prefer rules, since models never belong in a well-run product.
C. When today's logic can be written down completely and correctly as rules, use rules. Save model complexity for genuinely fuzzy, pattern-based judgment.
D. It doesn't matter which one you pick, as long as monitoring is in place.
Show hint
Look at the direct answer's first sentence.
Show answer
C. The kill criteria, can the logic be written down completely as rules, is what decides it, not a blanket rule against models or a blanket preference for them.
Short answer, apply it yourself
6. Think of a product you use that makes some decision by machine learning. Name one part of that decision that's probably fully known, stable logic a rules engine could handle instead.
Show hint
Ask which part of the decision has math or a policy behind it that a person could already write down today.
Show answer
Model answer: A ride-hailing app uses a model to set surge pricing, but whether a rider even qualifies for a promo code, a flat percentage off if their account is under 30 days old, is fully known logic. That check doesn't need a model at all, just a date comparison.
Before you close the answer
Why this works
Tests whether you can tell a fully known decision from a genuinely fuzzy one, instead of reaching for a model because it feels more modern. Most candidates either distrust every rules-based system as old-fashioned, or trust every model as automatically smarter. The real judgment is asking, case by case, whether the logic was ever actually unknown.
Follow-up traps
"Doesn't the model eventually learn to be stable, once it's trained on enough data?" Response: No, because the instability wasn't from bad training. It came from retraining itself, on a number that was never supposed to move. A new model version can always find a slightly different fit, even when the real answer hasn't changed at all.
"What if demand for the screws actually does shift someday? Won't the rules engine miss it?" Response: That's exactly what the kill test is for. Revisit a SKU when its real demand pattern actually starts changing, not on a fixed schedule. Checked against real sales once a quarter, a formula catches a genuine shift just as well as a model would, for a fraction of the cost.
If pressed
The formula's safety-stock term uses a 1.65 z-score, targeting a 95 percent service level. Push that to 99 percent for something Stagbrook genuinely can't afford to run out of, and the formula's answer moves in one predictable direction. You just change one input. No retraining, and no ten weeks of guessing required.
From U2xAI Academy
From answering questions to owning outcomes.
A live workshop where you ship a working AI agent, defend a launch decision, and walk away with a portfolio recruiters can't wave off, not just more questions to study.