Tokens and Tokenization in LLM
Tokenization in LLMs is the process of splitting text into tokens, the units a model reads and generates, and mapping each token to an integer ID from a fixed vocabulary. Tokens in LLMs are usually whole words, subword fragments or single characters, and their count determines context usage, latency and API cost.
- Token: The smallest unit a large language model processes, typically a common word, part of a rare word or a punctuation symbol.
- Vocabulary: The fixed set of tokens a tokenizer recognises, commonly tens of thousands to a few hundred thousand entries.
- Subword splitting: Rare words and long identifiers are broken into smaller known pieces, so no input is ever out of vocabulary.
- Token IDs: Each token maps to an integer that selects one row of the model's embedding table.
- Billing unit: Commercial APIs price requests and limit context in tokens rather than in characters or words.
For example, a tokenizer may split the Python identifier calculate_total into calc, ulate, _ and total, so a single function name consumes four tokens of the model's context.
Key Characteristics of Tokenization in LLMs
- Learned vocabulary: Algorithms such as byte pair encoding (BPE) build the vocabulary by repeatedly merging the most frequent adjacent symbol pairs in a training corpus.
- Model-specific behaviour: Each model family ships its own LLM tokenizer, so identical text yields different token counts across models.
- Language-dependent efficiency: English prose averages roughly four characters per token, while Hindi and other Indic scripts often require more tokens per word.
- Whitespace and case sensitivity: A leading space or a capital letter can produce a different token, so
Totalandtotalmay receive separate IDs. - Byte fallback: Byte-level tokenizers represent any Unicode character, including emoji and unusual symbols, as a sequence of byte tokens.
- Code-specific patterns: Indentation, brackets and camelCase identifiers are tokenized differently from prose, which affects how much source code fits in a prompt.
How Tokenization in LLMs Works
- Pre-tokenization: The text is first divided on whitespace and punctuation into rough word-like chunks.
- Subword matching: Each chunk is divided into the longest or most frequent pieces that exist in the vocabulary.
- Fallback handling: Characters without a matching piece become single-character or byte-level tokens.
- ID mapping: Each token is replaced by its integer ID, producing the sequence that enters the pipeline described in how LLMs work.
- Embedding lookup: The model converts each ID into a vector, as covered in embeddings in LLMs.
- Detokenization: Generated IDs are mapped back to text pieces and concatenated into the final output.
How Byte Pair Encoding Chooses Its Merges
Byte pair encoding (BPE) relies on counting rather than calculus: it counts how often two adjacent symbols occur across the corpus, weights each occurrence by word frequency, and merges the most frequent pair into a new vocabulary symbol.
- : the set of distinct words or identifiers in the training corpus.
- : the frequency of the word, meaning how many times it appears in the corpus.
- : the number of positions inside the word where one symbol is immediately followed by the other.
- : the pair with the highest weighted count, which becomes a single new symbol.
- : the size of the initial vocabulary, usually individual characters or the 256 possible byte values.
- : the number of completed merges, since every merge adds exactly one entry to the vocabulary.
In words, the training procedure repeats four operations:
- Initial split: Every word in the corpus begins as a sequence of individual characters.
- Pair counting: The count formula calculates a weighted frequency for every adjacent pair across the corpus.
- Merging: The most frequent pair is joined wherever it appears and is appended to the vocabulary.
- Repetition: Counting and merging continue until the vocabulary reaches its target size, and the ordered list of merges becomes the tokenizer.
At tokenization time, new text is split into characters, and the stored merges are applied in exactly the same order in which they were learned.
Byte Pair Encoding Example on Code Identifiers
The corpus contains six identifier fragments with their frequencies, and the initial vocabulary contains seven distinct characters.
- First merge: The pair of
eandtoccurs once in every identifier, so its weighted count is the largest. - Second merge: The new symbol
etnow followssin four identifiers, which makessetthe next vocabulary entry. - Third merge: The pair of
gandetoccurs ingetandgetter, producing the symbolget. - Tie-breaking: Two pairs share the fourth-highest count, and an alphabetical rule selects the pair of
eandr.
from collections import Counter
# Identifier fragments from a codebase, with how often each appears
corpus = {"get": 6, "set": 5, "reset": 3, "getter": 2, "setter": 2, "offset": 1}
# Start with every word split into single characters
words = {w: list(w) for w in corpus}
def pair_counts():
counts = Counter()
for w, symbols in words.items():
for a, b in zip(symbols, symbols[1:]):
counts[a, b] += corpus[w]
return counts
def merge(pair):
for w, s in words.items():
out, i = [], 0
while i < len(s):
if i < len(s) - 1 and (s[i], s[i + 1]) == pair:
out.append(s[i] + s[i + 1])
i += 2
else:
out.append(s[i])
i += 1
words[w] = out
vocab = sorted({c for w in corpus for c in w})
print("Start vocabulary:", len(vocab), "symbols")
for step in range(1, 5):
counts = pair_counts()
# Most frequent pair; ties go to the alphabetically first pair
best = min(counts, key=lambda p: (-counts[p], p))
merge(best)
vocab.append(best[0] + best[1])
print(f"Merge {step}: {best[0]} + {best[1]} -> {best[0] + best[1]} (count {counts[best]})")
print("Vocabulary size:", len(vocab))
for w in corpus:
print(f" {w:7} {words[w]}")Start vocabulary: 7 symbols
Merge 1: e + t -> et (count 19)
Merge 2: s + et -> set (count 11)
Merge 3: g + et -> get (count 8)
Merge 4: e + r -> er (count 4)
Vocabulary size: 11
get ['get']
set ['set']
reset ['r', 'e', 'set']
getter ['get', 't', 'er']
setter ['set', 't', 'er']
offset ['o', 'f', 'f', 'set']- Frequent fragments become single tokens:
getandsetrequire one token each after three merges, while the infrequentoffsetstill requires four. - Pieces are shared:
reset,setterandoffsetall reuseset, which is how production tokenizers reuse fragments such as_idacross many identifiers.
The number of merges determines the vocabulary size, so it balances a larger embedding table against shorter token sequences and a lower cost for every request.
Example: Subword Tokenization of Python Code
The program below implements a greedy longest-match tokenizer with a small vocabulary and compares character and token counts for several code fragments.
import re
# A small subword vocabulary; real tokenizers learn tens of thousands of pieces
VOCAB = {"calc", "ulate", "total", "get", "user", "by", "id", "_", "(", ")",
":", "def", "items", "item", "s", "response", "json", "."}
def split_word(word):
"""Greedy longest match: take the longest vocabulary piece at each position."""
pieces, i = [], 0
while i < len(word):
for j in range(len(word), i, -1):
if word[i:j].lower() in VOCAB:
pieces.append(word[i:j])
i = j
break
else: # unknown character: keep it as a single-character token
pieces.append(word[i])
i += 1
return pieces
def tokenize(text):
tokens = []
for word in re.findall(r"[A-Za-z]+|\d+|[^\sA-Za-z\d]", text):
tokens.extend(split_word(word))
return tokens
for text in ["def calculate_total(items):", "getUserById(id)", "response.json()", "qx9_zz"]:
toks = tokenize(text)
print(f"{text!r}: {len(text)} chars -> {len(toks)} tokens")
print(" ", toks)'def calculate_total(items):': 27 chars -> 9 tokens
['def', 'calc', 'ulate', '_', 'total', '(', 'items', ')', ':']
'getUserById(id)': 15 chars -> 7 tokens
['get', 'User', 'By', 'Id', '(', 'id', ')']
'response.json()': 15 chars -> 5 tokens
['response', '.', 'json', '(', ')']
'qx9_zz': 6 chars -> 6 tokens
['q', 'x', '9', '_', 'z', 'z']- Common names are compact:
response.json()needs only 5 tokens for 15 characters, because every piece already exists in the vocabulary. - Identifiers split into known pieces:
getUserByIdbecomesget,User,ByandId, because the vocabulary holds each piece but not the whole name. - Unfamiliar strings are expensive:
qx9_zzcontains no known pieces, so each of its 6 characters becomes a separate token.
Applications of Tokenization
- Cost estimation: Counting tokens before a request predicts the API spend of a code review pipeline.
- Context budgeting: Token counts decide how much of a repository or log fits in the context window.
- Prompt compression: Removing redundant whitespace and boilerplate reduces token usage per request.
- Rate limiting: Providers enforce throughput limits measured in tokens per minute.
- Chunking for retrieval: Documents are divided into token-sized chunks before indexing for RAG.
- Multilingual products: Measuring token counts per language reveals cost differences between English and Hindi users.
Advantages
- Open vocabulary: Subword splitting represents any word, identifier or typing error without an unknown-token failure.
- Compact sequences: Frequent words become single tokens, which shortens input compared with character-level processing.
- Shared structure: Related words such as
parse,parserandparsingcan share pieces, which helps the model generalise. - Deterministic output: A given tokenizer always converts identical text into the same token sequence.
Limitations
- Uneven cost: Non-English text and unusual identifiers consume more tokens for equivalent content.
- Character blindness: Models see tokens rather than letters, so counting characters or reversing strings is error-prone.
- Arithmetic difficulty: Numbers split into irregular pieces, which complicates digit-level reasoning.
- Tokenizer lock-in: A vocabulary cannot be changed after training without retraining the embedding layer.
- Formatting sensitivity: Minor whitespace changes can alter tokenization and slightly change model output.
Quick Quiz
Pick an answer to check yourself. Nothing is saved.
Question 1 / 3
1. What does byte pair encoding do when building a vocabulary?
Frequently Asked Questions
What is a token in LLM terminology?
A token is the basic unit of text that a language model reads and generates. It can be a whole word, part of a word, a punctuation mark or a single character.
How many words are in 1,000 tokens?
For typical English prose, 1,000 tokens correspond to roughly 700 to 800 words. Code, non-English text and unusual identifiers usually need more tokens per word, so the ratio varies by content and tokenizer.
Why do LLMs use subword tokens instead of whole words?
A whole-word vocabulary would need millions of entries and would still fail on new words and identifiers. Subword pieces keep the vocabulary manageable while representing any possible input.
Why does Hindi text use more tokens than English?
Most tokenizers are trained on corpora dominated by English, so English words often become single tokens. Hindi words are split into more pieces, which raises token counts and cost for the same content.
Related Articles
- What is a Large Language Model (LLM)Learn what a large language model is, how an LLM predicts the next token, its key characteristics, uses and limits, with a Python example and diagram.
- How LLMs WorkLearn how LLMs work from training to inference: tokenization, embeddings, attention layers and decoding, with a Python trace of one forward pass.
- Context Window in LLMLearn what the context window of an LLM is, how tokens fill it, what happens past the limit and how to fit long logs, with a Python truncation example.
- Embeddings in LLMLearn what embeddings in LLMs are, how text becomes vectors, how cosine similarity compares meaning and where embeddings are used, with Python code.
- Transformer Architecture ExplainedTransformer architecture explained: self-attention, query, key and value vectors, multi-head attention and stacked layers, with Python attention code.
- Temperature and Top-p in LLMLearn how temperature and top-p control LLM sampling, how scaling and filtering change next-token odds, with a Python code example and a clear diagram.