When do you use rule-based infra vs. NN vs. LLM models?
Transcript
Read the full transcript (1,383 words)
[INTERVIEWER] When do you use rule-based infra vs. NN vs. LLM models? When do you use rules versus a trained model versus an LLM? The answer that quietly fails is reaching for the LLM on everything because it's the shiny one. The real answer is to match the tool to the problem and reach for the simplest thing that clears the bar.
Every step up in complexity costs you latency, money, and predictability. The order you reach in is rules first, then a trained model, then an LLM. And the best systems don't pick just one. They layer all three. Let me walk you through it. This tests engineering judgement, specifically whether you'd over-engineer. A PM who puts an LLM behind every button ships something slow, expensive, and non-deterministic where it didn't need to be.
The signal they want is that you justify complexity by need rather than novelty. So here's the ladder. The bottom rung is rules. Deterministic logic like if-then statements, regex, and thresholds. You use rules when the logic is known, fixed, and has to be exact and auditable. Hard compliance limits. Business rules. Input validation. Guardrails. Hard safety blocks. Anywhere a wrong answer is genuinely unacceptable and you need a guarantee, not a probability.
Their strengths are that they're deterministic, transparent, free, instant, and fully testable. You can prove exactly what they'll do. Their weaknesses are that they're brittle, they don't generalise at all, and the rule set rots as it grows. Every new edge case is another line someone has to maintain. But for anything that must be exact, rules aren't the old-fashioned choice.
They're the correct one. The middle rung covers neural nets and classic ML. These are trained models like gradient boosting, CNNs, and smaller fine-tuned networks. You reach for these when you've got labelled data and a well-defined, repeated prediction with a fixed output space. Fraud scoring. Ranking and recommendations. Churn prediction. Image classification. Demand forecasting. Spam detection. Their strengths are that they're cheap and fast at inference, they hit high accuracy on the exact task they were trained for, and you can calibrate their confidence.
Their weaknesses are that they need labelled data and a training pipeline to exist in the first place, and they only do the one narrow task they were trained on. Ask a fraud model to summarise a document and you get nothing. But for a high-volume, repeated prediction, nothing beats a small trained model on cost and speed. The top rung is the LLM.
You use it for open-ended natural language and reasoning over a long tail of inputs, covering cases where you have few labels and you need real flexibility. Summarisation. Open question-answering. Extraction from messy, unstructured text. Chat. Code. Agentic tool use. Its strengths are that it's general, you can prototype with a prompt and no training at all, and it handles inputs it's never seen.
The weaknesses are the ones you must respect. It's expensive, it has higher latency, it's non-deterministic so the same input can give different answers, it can hallucinate, and it's hard to fully test. So the LLM is powerful, but it's the top of the ladder for a reason. You climb to it when the lower rungs genuinely can't do the job.
Here's the judgement that actually scores, which weak answers miss entirely. These three aren't competitors you pick between. They compose. Real systems layer them. Rules act as guardrails wrapped around an LLM. A cheap classifier routes traffic, sending things to the LLM only when they actually need it. The LLM handles the messy long tail that a classifier can't. So you don't ask which one, you ask which one for which part.
And the discipline that goes with it is clear. Don't reach for an LLM where a regex or a gradient-boosted tree is cheaper, faster, and more reliable. That restraint is the senior signal. Let's state the tradeoff plainly. It's cost and determinism versus flexibility. A rule or a small classifier is roughly a thousand times cheaper than an LLM call, and it runs in under a millisecond.
An LLM costs cents per call and takes hundreds of milliseconds. So you pay that premium in money and latency only where the flexibility is genuinely worth it. On the easy, structured, high-volume stuff, paying LLM prices is just burning money for no gain. Let me make it concrete with a moderation and support pipeline, because you can see all three layered in one flow.
Input validation and PII redaction, stripping out personal data before anything else touches it, that's rules. It's deterministic and auditable, exactly what you want on something legally sensitive. Intent classification, spam detection, and priority scoring, that's a fine-tuned small classifier. A BERT-class model that runs in under ten milliseconds and costs almost nothing per call. Drafting a reply to a novel, free-text customer request, that's the LLM, because you genuinely can't enumerate the inputs.
And a hard policy block like never give legal advice, back to rules, sitting on the output this time. Now, on moderation specifically, a regex blocklist plus a trained toxicity classifier catch about ninety-five percent of the clear cases cheaply and deterministically. Only the ambiguous five percent goes to an LLM for the nuanced judgement call. Compare that to running the LLM on all the traffic.
That would be roughly twenty times the cost, and slower, for zero accuracy gain on the easy cases the classifier already nailed. Same quality, a fraction of the cost, because you matched each tool to its part of the job. A follow-up worth preparing for is when you built an LLM feature and it's now too slow and too expensive at scale.
What do you do? The answer is that you push work down the ladder. You look at your traffic and you find the slice that's actually repetitive and structured, and you pull it off the LLM onto a small classifier or a cache. If forty percent of your queries are near duplicates, a semantic cache handles those for almost nothing.
If a chunk of the traffic is really just classification dressed up as a chat, a fine-tuned small model does it in ten milliseconds. The LLM then only sees the genuine long tail, which is what it was for anyway. So the migration path is almost always the same direction. Start on the LLM to learn the problem fast, then demote the parts that turn out to be predictable.
That shows you understand the ladder isn't just a design-time choice. It's how you drive cost down after launch. What makes them lean in? First, you reach for the simplest tool first and justify every step up in complexity by need, the anti-over-engineering signal they're hunting for. Second, you show the three composing inside one system rather than giving a single-tool answer, so it's clear you've architected something real.
Third, you put actual cost and latency numbers on the difference. A thousand times cheaper, under a millisecond, twenty times the cost. This makes the tradeoff concrete instead of hand-wavy. Numbers are what turn a decent answer into a memorable one. The traps here match the exact failure modes of the AI hype cycle. Trap one is reaching for an LLM by default, including in the places where a rule or a small classifier is strictly better on every axis.
Trap two is treating the three as competitors you have to choose between, when the whole point is that they layer. Trap three is forgetting that rules are the right call, not the outdated one, for anything that has to be exact and auditable. If you put an LLM on a hard compliance limit, you've misunderstood the job. To recap.
Rules for logic that's known, fixed, and must be exact and auditable. Trained models for repeated, structured prediction where you have labels. LLMs for the open-ended, unstructured long tail where you need flexibility. Reach for the simplest tool that clears the bar, put cost and latency numbers on the jump, and remember the real systems layer all three in one pipeline.
The one line to carry in is that rules are for exact and auditable logic, trained models are for repeated structured prediction, LLMs are for the open-ended long tail, and real systems layer all three.