Small Language Models (SLM)
Small language models (SLMs) are language models with far fewer parameters than frontier large language models, typically a few billion or fewer. They trade some general capability for lower memory use, lower latency and lower cost, which lets them run on laptops, phones and modest servers.
- Compact size: It has fewer parameters, so its weights occupy far less memory than a large language model.
- Local inference: It can run as a local LLM on a developer laptop, a smartphone or an edge device without any network connection.
- Specialised strength: It performs well on focused tasks such as classification, information extraction and short code completions.
- Common families: Examples include Microsoft Phi, Google Gemma and the smaller variants of Llama and Qwen.
- Open weights: Most SLMs are released as open-weight models that organisations can download, modify and host independently.
For example, a code editor can run a 3-billion-parameter model locally to suggest the next line of Python without sending source code to a remote server.
Key Characteristics of Small Language Models
- No fixed threshold: There is no official parameter count that separates small from large models, and the boundary shifts over time.
- Quantization support: Weights are often stored at 8-bit or 4-bit precision, which reduces memory with a modest loss in quality.
- Knowledge distillation: Many SLMs are trained partly on outputs of larger models, a technique in which a large teacher model guides a smaller student.
- Low latency: Fewer calculations per token produce faster responses, which suits interactive code completion.
- Limited world knowledge: Fewer parameters store fewer facts, so SLMs depend more heavily on information supplied in the prompt.
- Easy customisation: Their small size makes fine-tuning faster and considerably cheaper than for large models.
How Small Language Models Work
- Training: The model is trained with the same next-token objective as larger models, usually on carefully filtered, high-quality data.
- Distillation: Outputs from a larger model are frequently used as additional training data to transfer specific capabilities.
- Compression: Quantization reduces the precision of each weight, and pruning removes weights that contribute little.
- Packaging: The compressed weights are packaged in a format that local runtimes such as llama.cpp or Ollama can load.
- Local inference: The runtime loads the weights into CPU, GPU or mobile memory and generates tokens directly on the device.
- Routing: Applications frequently send simple requests to the SLM and escalate complicated requests to a larger hosted model.
Quantization Formula and Memory Estimates
Two formulas explain why small language models fit on consumer devices. The first estimates the memory required for the weights, and the second, the quantization formula, describes how each weight is represented with fewer bits.
The memory required for the weights equals the number of parameters multiplied by the storage that each individual parameter occupies:
- : the total number of parameters in the model.
- : the numerical precision in bits per parameter, such as 16, 8 or 4, where dividing by 8 converts bits into bytes.
For an illustrative 3-billion-parameter model, 16-bit weights require bytes, approximately 6 GB, whereas 4-bit weights require bytes, approximately 1.5 GB. The Python example in the following section applies this calculation to four different model sizes.
Quantization maps each real-valued weight to a small integer code, and dequantization converts that code back into an approximate weight during inference:
The scale and the zero point are calculated from the range of the original weights and the range of the available integer codes:
- : an original weight, typically stored as a 16-bit or 32-bit floating-point number.
- : the corresponding integer code, where 4 bits allow sixteen codes from 0 to 15, so and .
- : the scale, which is the distance between two neighbouring codes measured in weight units.
- : the zero point, which is the integer code that represents the weight 0.0 exactly.
- : the dequantized weight used during inference, which is close to the original value but not always identical.
Worked example. The following steps quantize five illustrative weights, and , to 4-bit integer codes.
- Scale: , the uniform spacing between neighbouring integer codes.
- Zero point: , so the weight 0.00 is represented exactly by code 6.
- Quantize 0.34: , the integer representation of this particular weight.
- Dequantize: , which introduces an approximation error of 0.04.
- Error bound: Rounding displaces any weight inside the representable range by at most half a step, which equals here.
The program below applies both formulas to all five weights and reports the approximation error for each one.
# 4-bit asymmetric quantization: q = round(x / s) + z, dequantized x_hat = s * (q - z).
weights = [-0.60, -0.23, 0.00, 0.34, 0.90]
q_min, q_max = 0, 15 # 4 bits give 16 integer levels
s = (max(weights) - min(weights)) / (q_max - q_min) # scale: step between levels
z = round(-min(weights) / s) # zero point: the integer for 0.0
print(f"scale s = {s:.2f}, zero point z = {z}")
print(f"{'x':>6}{'q':>4}{'x_hat':>7}{'error':>7}")
for x in weights:
q = min(q_max, max(q_min, round(x / s) + z)) # clamp to the 4-bit range
x_hat = s * (q - z)
print(f"{x:>6.2f}{q:>4}{x_hat:>7.2f}{abs(x - x_hat):>7.2f}")
# Stored size of these 5 weights: 16-bit floats vs 4-bit integers
print(f"16-bit: {len(weights) * 16} bits, 4-bit: {len(weights) * 4} bits plus one s and z")scale s = 0.10, zero point z = 6
x q x_hat error
-0.60 0 -0.60 0.00
-0.23 4 -0.20 0.03
0.00 6 0.00 0.00
0.34 9 0.30 0.04
0.90 15 0.90 0.00
16-bit: 80 bits, 4-bit: 20 bits plus one s and z- Four times smaller: Each weight's storage decreases from 16 bits to 4 bits, and the entire group requires only one additional scale and zero point.
- Outliers widen the step: A single unusually large weight stretches the range, which increases the scale and therefore the rounding error for every other weight in the group.
- Groups in practice: Production formats store a separate scale and zero point for each small block of weights, so an outlier affects only its own block.
In practice, the precision per parameter determines which device can hold a model, and the accumulated rounding error determines how much output quality that memory saving costs.
Example: Estimating Small Language Model Memory in Python
The program below estimates the memory needed to hold model weights at different precisions and checks which models fit on a device.
# Memory needed just to hold a model's weights: parameters x bytes per parameter.
# Sizes are illustrative; real use also needs memory for the context (KV cache).
sizes = {"1B": 1e9, "3B": 3e9, "8B": 8e9, "70B": 70e9}
precisions = {"16-bit": 2.0, "8-bit": 1.0, "4-bit": 0.5}
device_gb = {"phone": 6, "laptop": 16}
print(f"{'model':6}" + "".join(f"{p:>9}" for p in precisions))
for name, params in sizes.items():
cells = "".join(f"{params * b / 1e9:>7.1f}GB" for b in precisions.values())
print(f"{name:6}{cells}")
# Which models fit in half of each device's memory at 4-bit?
for device, gb in device_gb.items():
fits = [n for n, p in sizes.items() if p * 0.5 / 1e9 <= gb / 2]
print(f"Fits on a {gb} GB {device} at 4-bit: {', '.join(fits)}")model 16-bit 8-bit 4-bit
1B 2.0GB 1.0GB 0.5GB
3B 6.0GB 3.0GB 1.5GB
8B 16.0GB 8.0GB 4.0GB
70B 140.0GB 70.0GB 35.0GB
Fits on a 6 GB phone at 4-bit: 1B, 3B
Fits on a 16 GB laptop at 4-bit: 1B, 3B, 8B- Precision matters: Converting weights from 16-bit to 4-bit precision reduces the memory requirement of every model by a factor of four.
- Size decides placement: At 4-bit, a 3B model needs about 1.5 GB and fits on a phone, while a 70B model needs about 35 GB and requires a server GPU.
- Half-memory rule: The calculation reserves half of each device's memory for the operating system, concurrent applications and the context cache.
Applications of Small Language Models
- Local code completion: Short inline suggestions inside an editor with low latency and no data leaving the machine.
- On-device assistants: An on-device LLM powers smartphone and laptop features that must operate without connectivity.
- Classification and routing: Sorting tickets, logs or requests before a larger model handles the difficult ones.
- Structured extraction: Extracting fields from error messages or bug reports into structured JSON output.
- Edge and embedded systems: Inference on industrial equipment, vehicles or network appliances.
- Privacy-sensitive workloads: Processing confidential source code or regulated data entirely inside the organisation.
Advantages
- Lower cost: Inference requires less hardware, so each request costs considerably less than on a large model.
- Lower latency: Fewer calculations per token produce faster responses for interactive tools.
- Data privacy: Local execution keeps source code and user data on the device, a key point in open source vs closed source LLMs.
- Offline availability: The model continues operating without an internet connection or external dependency.
- Faster iteration: Fine-tuning and evaluation cycles complete in hours rather than days.
Limitations
- Weaker reasoning: Multi-step reasoning and complex code generation are considerably less reliable than with frontier models.
- Less knowledge: Rare libraries and specialised facts are often missing, which increases LLM hallucination.
- Smaller context: Many SLMs support shorter context windows than large hosted models.
- Quantization loss: Aggressive compression can reduce output quality in ways that require testing to detect.
- Operational work: Local deployment requires packaging, updates and monitoring across many devices, as discussed in popular LLMs compared.
Quick Quiz
Pick an answer to check yourself. Nothing is saved.
Question 1 / 3
1. What is the main trade-off of a small language model?
Frequently Asked Questions
SLM vs LLM: what is the difference?
An SLM has far fewer parameters than an LLM, so it needs less memory and responds faster. An LLM usually handles complex reasoning and broad knowledge better.
Can a small language model run on a laptop?
Yes, many small language models run on a standard laptop, especially when their weights are quantized to 4-bit or 8-bit precision. Available memory decides which model sizes fit.
Are small language models good for coding?
They work well for short completions, simple refactoring and code explanation. Complex multi-file changes and difficult debugging usually need a larger model.
What is quantization in small language models?
Quantization stores each weight with fewer bits, for example 4 bits instead of 16. It reduces memory and speeds up inference, with a small loss in output quality.
Related Articles
- What is a Large Language Model (LLM)Learn what a large language model is, how an LLM predicts the next token, its key characteristics, uses and limits, with a Python example and diagram.
- Fine-Tuning LLMsLearn what fine-tuning LLM models means, how LoRA and full tuning work, when to use them, with a Python example that prepares chat-format training data.
- Open Source vs Closed Source LLMsCompare open source vs closed source LLM options on data control, cost, customisation and setup, with a comparison table and a pull request example.
- Popular LLMs (Claude, GPT, Gemini, Llama) ComparedCompare popular LLMs Claude, GPT, Gemini and Llama on weights, access, strengths and self-hosting, with a comparison table and a stack trace example.
- Tokens and Tokenization in LLMLearn how tokenization in LLMs splits text and code into subword tokens and why token counts drive cost and context, with a Python tokenizer example.
- How LLMs WorkLearn how LLMs work from training to inference: tokenization, embeddings, attention layers and decoding, with a Python trace of one forward pass.