…
Skip to content
Topics
On this page

Embeddings in LLM

Embeddings in LLMs are dense numerical vectors that represent tokens, sentences or documents, positioned so that items with similar meaning lie close together in vector space. A large language model uses token embeddings as the input to its layers, and dedicated embedding models produce vectors for semantic search and retrieval.

  • Vector: An ordered list of numbers, typically hundreds to a few thousand dimensions long.
  • Learned representation: Embedding values are adjusted during training, so tokens used in similar contexts converge towards similar vectors.
  • Token embeddings: The first layer of an LLM maps each token ID to a vector before any attention is applied.
  • Text embeddings: Dedicated embedding models compress a whole sentence, function or document into a single vector, and a vector database stores these vector embeddings for retrieval.
  • Similarity measure: Cosine similarity compares the direction of two vectors, where 1 indicates identical direction and 0 indicates no relationship.
Code snippets and a query as embeddings in vector spaceA two-dimensional sketch of an embedding space. The query "read a text file line by line" sits inside a cluster with two file-reading snippets, f.readlines() and open(...).read(). Snippets for HTTP requests, summing prices and handling a KeyError sit far from the query. Real embeddings have hundreds of dimensions; these positions are illustrative.File reading clusterf.readlines()open(...).read()requests.get(url)sum(prices)except KeyErrorStar: query "read a text file line by line"; close points mean similar meaning
Code snippets and a query as embeddings in vector space

For example, a code search tool embeds with open(path) as f: lines = f.readlines() and the query "read a text file line by line" into nearby vectors, although the two share almost no wording.

Key Characteristics of Embeddings in LLMs

  • Dense representation: Every dimension holds a real number, unlike sparse one-hot vectors that contain a single 1 among zeros.
  • Semantic geometry: Distance and direction encode relationships, so related concepts form clusters in the vector space.
  • Contextual refinement: Inside a transformer, attention updates each token vector, so the word bank in two different sentences ends with different representations.
  • Fixed dimensionality: All vectors from one model have the same length, which allows direct mathematical comparison.
  • Model-specific spaces: Vectors from different embedding models occupy incompatible spaces and cannot be compared with each other.
  • Cross-language alignment: Multilingual and code-aware models place equivalent content from different languages near each other.

How Embeddings in LLMs Work

  1. Tokenization: Input text is split into tokens and integer IDs, as explained in tokenization in LLMs.
  2. Lookup: Each ID selects one row of a learned embedding matrix, producing one vector per token.
  3. Contextualisation: Transformer layers exchange information between token vectors through attention, as described in transformer architecture.
  4. Pooling: An embedding model combines the token vectors, for example by averaging them, into one vector for the entire text.
  5. Normalisation: The vector is often scaled to unit length, so dot product and cosine similarity produce the same ranking.
  6. Comparison: Applications compute the similarity between a query vector and stored vectors to identify the closest matches.

Computing Cosine Similarity

Embedding similarity is computed with three related formulas, namely the dot product, the vector norm and the cosine similarity, which combines the first two into a single normalised score.

The dot product multiplies the corresponding entries of two vectors and then adds all of those products together:

The norm measures the length of a vector, and it generalises the Pythagorean theorem from two dimensions to dimensions:

The cosine similarity formula divides the dot product by both norms, which removes the effect of vector length and leaves only the angle between the two vectors:

  • : two embedding vectors, such as a natural-language query and a stored code snippet.
  • : the individual component in position of each vector.
  • : the number of dimensions, which equals 3 in the worked example and several hundred in production models.
  • : the norm, or Euclidean length, of the vector .
  • : the angle between the vectors, whose cosine equals 1 for identical directions, 0 for perpendicular directions and -1 for opposite directions.

Normalisation divides every component by the vector's norm, so the resulting unit vector has a length of exactly 1, and for two unit vectors the ordinary dot product already equals the cosine similarity:

Worked example. Suppose the three dimensions represent file operations, text processing and network communication. The query "read a text file" has the vector , and the snippet open(path).read() has .

  1. Dot product: .
  2. Norm of the query: .
  3. Norm of the snippet: .
  4. Cosine similarity: , which corresponds to an angle of approximately 17.8 degrees.
  5. A longer vector: The snippet requests.get(url).text has , which produces a larger dot product of 30, but its norm of 15 reduces the cosine to .

The program below repeats these calculations for a code snippet, a documentation page and an HTTP request, and then verifies the normalised version.

Python
import math

# 3-dimensional embeddings. Dimensions: file I/O, text handling, network.
query = [2, 2, 1]  # "read a text file"
items = {
    "open(path).read()": [6, 3, 2],
    "docs: File handling guide": [2, 1, 2],
    "requests.get(url).text": [0, 9, 12],
}

def dot(a, b):
    return sum(x * y for x, y in zip(a, b))

def norm(a):
    return math.sqrt(dot(a, a))

def normalise(a):
    n = norm(a)
    return [x / n for x in a]

print(f"|query| = {norm(query):.0f}")
print(f"{'item':27}{'dot':>5}{'|b|':>5}{'cosine':>8}{'angle':>7}")
for name, vec in items.items():
    cos = dot(query, vec) / (norm(query) * norm(vec))
    angle = math.degrees(math.acos(cos))
    print(f"{name:27}{dot(query, vec):>5}{norm(vec):>5.0f}{cos:>8.3f}{angle:>6.1f}°")

# After normalising, the dot product equals the cosine similarity.
u, v = normalise(query), normalise(items["open(path).read()"])
print(f"unit vectors: dot = {dot(u, v):.3f}, lengths = {norm(u):.1f} and {norm(v):.1f}")
Output
|query| = 3
item                         dot  |b|  cosine  angle
open(path).read()             20    7   0.952  17.8°
docs: File handling guide      8    3   0.889  27.3°
requests.get(url).text        30   15   0.667  48.2°
unit vectors: dot = 0.952, lengths = 1.0 and 1.0
Cosine similarity ranks vectors by their angle to the query, not by their lengthThe query vector [2, 2, 1] with length 3 lies along the bottom. Three item vectors start at the same origin, drawn at their true angle to the query and at their true length. open(path).read() has length 7, an angle of 17.8 degrees and cosine 0.952. A docs guide has length 3, an angle of 27.3 degrees and cosine 0.889. requests.get(url).text has length 15, the largest dot product of 30, but the widest angle of 48.2 degrees and the lowest cosine, 0.667. The vectors are three-dimensional, so only their angle to the query is shown faithfully. The numbers are illustrative and match the worked example.query [2, 2, 1]1231. open(path).read()17.8°, cosine 0.9522. docs: File handling guide27.3°, cosine 0.8893. requests.get(url).text48.2°, cosine 0.6673 has the largest dot product (30)
Cosine similarity ranks vectors by their angle to the query, not by their length
  • Length distorts the dot product: The HTTP snippet receives the largest raw dot product only because its vector is long, whereas cosine similarity correctly ranks it last.
  • Normalised vectors agree: After normalisation to unit length, the dot product returns exactly the same value, 0.952, as the cosine similarity.
  • Dot product vs cosine similarity: The two measures produce identical rankings only when every stored vector has the same length.

In practice, embeddings are usually normalised once during indexing, so a vector database can rank millions of items with an inexpensive dot product instead of recomputing norms for every comparison.

The program below ranks five Python snippets against a natural-language query using hand-made four-dimensional embeddings and cosine similarity.

Python
import math

# Hand-made 4-dimensional embeddings for code snippets. The dimensions stand
# for: file I/O, HTTP requests, summing numbers, error handling.
snippets = {
    "open('data.csv').read()": [0.9, 0.1, 0.0, 0.1],
    "with open(path) as f: lines = f.readlines()": [0.8, 0.0, 0.1, 0.3],
    "requests.get(url).json()": [0.1, 0.9, 0.0, 0.2],
    "total = sum(prices)": [0.0, 0.0, 0.9, 0.0],
    "try: ... except KeyError: ...": [0.1, 0.1, 0.0, 0.9],
}

def cosine(a, b):
    dot = sum(x * y for x, y in zip(a, b))
    norm = math.sqrt(sum(x * x for x in a)) * math.sqrt(sum(y * y for y in b))
    return dot / norm

query = "read a text file line by line"
query_vec = [0.85, 0.0, 0.05, 0.2]  # the embedding a model might give the query

print(f"Query: {query}")
ranked = sorted(snippets.items(), key=lambda kv: -cosine(query_vec, kv[1]))
for code, vec in ranked:
    print(f"  {cosine(query_vec, vec):.3f}  {code}")
Output
Query: read a text file line by line
  0.990  with open(path) as f: lines = f.readlines()
  0.985  open('data.csv').read()
  0.333  try: ... except KeyError: ...
  0.154  requests.get(url).json()
  0.057  total = sum(prices)
  • Meaning over wording: Both file-reading snippets score above 0.98, because their vectors point in almost the same direction as the query.
  • Unrelated code scores low: total = sum(prices) scores only 0.057, because its vector emphasises a different dimension.
  • Real embeddings are learned: A trained model produces hundreds of dimensions without human-readable meanings, and a vector database stores them for fast search.

Applications of Embeddings

  • Semantic code search: Finding functions by describing their behaviour, a technique known as semantic search.
  • Retrieval-augmented generation: Selecting relevant documentation passages for RAG.
  • Duplicate detection: Identifying near-identical bug reports or error messages.
  • Clustering: Grouping log lines or support tickets by topic.
  • Recommendation: Suggesting related documentation pages or code examples.
  • Classification: Training lightweight classifiers on embeddings instead of raw text.

Advantages

  • Captures meaning: Similar concepts match even when they use different vocabulary.
  • Efficient comparison: Similarity reduces to fast vector arithmetic over millions of stored items.
  • Reusability: One set of embeddings supports search, clustering and classification.
  • Cross-lingual matching: Multilingual models can match an English query to a Hindi document.

Limitations

  • Opaque dimensions: Individual dimensions rarely correspond to interpretable concepts.
  • Model lock-in: Changing the embedding model requires re-embedding every stored item.
  • Weak on exact terms: Identifiers, error codes and version numbers may match poorly, so keyword search is often combined with vector search.
  • Storage cost: Millions of high-dimensional vectors require substantial memory and specialised indexes.
  • Inherited bias: Embeddings reflect associations, including social biases, present in the training data.

Quick Quiz

Pick an answer to check yourself. Nothing is saved.

Question 1 / 3

  1. 1. What does a cosine similarity close to 1 indicate?

Frequently Asked Questions

What are embeddings in simple terms?

Embeddings are lists of numbers that represent the meaning of pieces of text. Texts with similar meaning receive similar lists, so a computer can compare meaning with arithmetic.

What is the difference between token embeddings and sentence embeddings?

A token embedding represents one token and is the input to an LLM's layers. A sentence embedding represents a whole passage as one vector and is produced by an embedding model for search and retrieval.

Why is cosine similarity used to compare embeddings?

Cosine similarity measures the angle between two vectors and ignores their length. This makes it a stable measure of how closely two pieces of text point in the same semantic direction.

Can embeddings from different models be compared?

No. Each model learns its own vector space, so vectors from two different models are not aligned. Every stored item must be embedded with the same model as the query.