What is a Large Language Model (LLM)
A large language model (LLM) is a neural network trained on very large collections of text and code to predict the next token in a sequence. By repeating this prediction and feeding each generated token back into its input, it can answer questions, summarise documents, translate text and generate code from natural-language instructions.
- Next-token prediction: It generates output one token at a time, where a token is a word, a subword fragment or a punctuation symbol.
- Scale: It is trained on enormous text corpora and stores what it learns in billions of parameters, the numerical weights adjusted during training.
- Transformer-based: Most modern LLMs use the transformer architecture, a neural network design built around a mechanism called attention.
- General purpose: A single model performs many language tasks without separate training for each one.
- Generative: It is a category of generative AI that composes new text instead of retrieving stored answers.
- LLM examples: ChatGPT, Claude and Gemini are conversational products built on LLMs.
For example, given def calculate_total(items): followed by total =, an LLM assigns probabilities to candidate tokens such as 0, sum( and [], selects the most probable one and repeats the process.
Key Characteristics of Large Language Models
- Pre-training on large corpora: It learns grammar, factual associations and programming conventions from books, websites, documentation and source code.
- Probabilistic output: It computes a probability distribution over its entire vocabulary at each step, so an identical prompt can produce different responses.
- Limited context: It processes only a fixed number of tokens per request, a limit called the context window.
- Prompt-driven: Its behaviour is controlled through natural-language instructions, a discipline known as prompt engineering.
- Adaptable: It can be specialised for a domain through fine-tuning or by supplying relevant documents with each request.
- Static knowledge: Its knowledge ends at the training cut-off, the date on which its training data was collected.
How a Large Language Model Works
The steps below describe inference, the stage in which a trained model generates text. A detailed walkthrough, including training, appears in how LLMs work.
- Tokenization: A tokenizer splits the input into tokens and maps each one to an integer identifier from the model's fixed vocabulary.
- Embedding: Each identifier is converted into an embedding, a numerical vector that represents meaning, so tokens used in similar contexts receive similar vectors.
- Attention: Each token computes weights over every earlier token to determine which ones are most relevant, which lets it resolve a pronoun to the correct noun.
- Scoring: The final layer produces a probability for every token in the vocabulary, and tokens consistent with the context receive the highest values.
- Sampling: The decoder selects one token from this distribution, and settings such as temperature control how often less probable tokens are chosen.
- Repetition: The selected token is appended to the input, and the loop continues until the model emits a stop token or reaches a length limit.
Example: Next-Word Prediction in Python
The program below counts which token follows which across nine lines of Python code, then uses those frequencies to complete a line.
import re
from collections import Counter, defaultdict
# Training text: short lines of Python code
lines = [
"total = 0", "count = 0", "total = sum(prices)",
"names = []", "for item in items:", "total = total + item.price",
"return total", "return total", "return names",
]
# Split each line into tokens and count which token follows which
following = defaultdict(Counter)
for line in lines:
tokens = re.findall(r"\w+|[^\w\s]", line)
for current, nxt in zip(tokens, tokens[1:]):
following[current][nxt] += 1
def predict(token):
counts = following[token]
n = sum(counts.values())
return [(t, round(c / n, 2)) for t, c in counts.most_common()]
print("After '=':", predict("="))
print("After 'return':", predict("return"))
# Greedy completion: add the most likely token until none is known
tokens = ["total"]
while predict(tokens[-1]) and len(tokens) < 6:
tokens.append(predict(tokens[-1])[0][0])
print("Completion:", " ".join(tokens))After '=': [('0', 0.4), ('sum', 0.2), ('[', 0.2), ('total', 0.2)]
After 'return': [('total', 0.67), ('names', 0.33)]
Completion: total = 0- Probabilities from frequencies: After
=, the token0scores 0.4 because it followed=in two of five occurrences. An LLM computes comparable probabilities from learned parameters instead of raw counts. - Greedy decoding: Selecting the highest-probability token at each step extends
totalintototal = 0, the same iterative loop an LLM performs during code completion. - Why real LLMs need more: This model conditions on only one previous token, so it ignores the function name and the
itemsargument. Production LLMs learn from billions of lines of code and apply attention across thousands of earlier tokens.
Applications of Large Language Models
- Code completion: Suggesting subsequent lines of code inside an editor while a developer types.
- Summarisation: Condensing incident reports and pull request discussions into concise summaries.
- Customer support: Drafting responses to customer queries for human agents to review.
- Translation: Converting text between languages, such as English and Hindi.
- Question answering: Answering from internal documentation through RAG, which retrieves relevant passages before generation.
- Agentic systems: Serving as the reasoning component in agentic AI systems that plan steps and invoke tools.
Advantages
- Versatility: A single model handles classification, drafting, summarisation and translation.
- Natural-language interface: Users specify tasks in ordinary language instead of code or fixed commands.
- Speed: It drafts text in seconds that would take a person minutes.
- Adaptability: It can be customised for an organisation through fine-tuning or retrieved documents.
- Scalability: One deployed model can serve many concurrent users.
Limitations
- Hallucination: It can state incorrect information with apparent confidence, a failure known as LLM hallucination.
- Knowledge cut-off: It knows nothing about events after its training date unless they are supplied in the prompt.
- Cost: Training requires thousands of specialised accelerators, and every inference request also incurs a computational cost.
- Bias: It can reproduce biases present in the human-written text used for training.
- Context limit: Text beyond the context window is not processed, so long inputs must be truncated or split.
Quick Quiz
Pick an answer to check yourself. Nothing is saved.
Question 1 / 3
1. What is the core task a large language model is trained to do?
Frequently Asked Questions
What is LLM in AI, and what is its full form?
The LLM full form is Large Language Model. It is an AI model trained on large amounts of text to understand and generate language. The word "large" refers to both the training data and the number of parameters.
Is ChatGPT an LLM?
ChatGPT is a chat product, and the engine behind it is a large language model. The product adds a chat interface, safety filters and other features around the model.
What is the difference between an LLM and a search engine?
A search engine retrieves existing pages that match a query. An LLM generates new text from patterns learned during training, so its output is not looked up and can be wrong.
Why do LLMs give wrong answers?
An LLM predicts likely text, not verified facts, so it can produce fluent but false statements. This is called hallucination. Grounding the model in trusted documents and reviewing important outputs reduces the risk.
What is the difference between training and inference?
Training is the stage in which the model adjusts its parameters by predicting hidden tokens in large text collections. Inference is the stage in which the trained model generates new text for a prompt.
Related Articles
- How LLMs WorkLearn how LLMs work from training to inference: tokenization, embeddings, attention layers and decoding, with a Python trace of one forward pass.
- Tokens and Tokenization in LLMLearn how tokenization in LLMs splits text and code into subword tokens and why token counts drive cost and context, with a Python tokenizer example.
- Transformer Architecture ExplainedTransformer architecture explained: self-attention, query, key and value vectors, multi-head attention and stacked layers, with Python attention code.
- LLM Hallucination: Causes and FixesLearn what LLM hallucination is, why models invent facts and code APIs, and how to reduce it with RAG and validation, with a runnable Python checker.
- Popular LLMs (Claude, GPT, Gemini, Llama) ComparedCompare popular LLMs Claude, GPT, Gemini and Llama on weights, access, strengths and self-hosting, with a comparison table and a stack trace example.
- What is Generative AILearn what generative AI is, how it creates text, code and images from learned patterns, its traits, uses and limits, with a Python example and diagram.