Memory in AI Agents
Memory in AI agents is the set of mechanisms that preserve information across model requests, because a language model retains nothing between separate invocations. Short-term memory holds the recent conversation inside the context window, while long-term memory stores durable knowledge in an external database and retrieves only the relevant records when the agent requires them.
- Short-term memory: Also called conversation memory, it holds the recent messages, tool calls and tool results that are resent to the model with every request.
- Long-term memory: Persistent records, such as resolved incidents or user preferences, kept in a vector store or a key-value store.
- Summarisation: Older conversation turns are condensed into a brief summary, so the essential facts survive after the raw messages are removed.
- Retrieval: A search step selects the stored records most relevant to the current task and inserts them into the prompt.
- Write policy: A rule that determines which observations are important enough to persist for future sessions.
For example, a DevOps agent investigating checkout-api timeouts keeps the latest log excerpts in short-term memory and recalls a similar resolved incident from long-term memory.
Key Characteristics of Memory in AI Agents
- Stateless model: The model itself remembers nothing, so the application reconstructs the relevant context for every request.
- Bounded capacity: Short-term memory is limited by the context window, measured in tokens.
- External persistence: Long-term memory survives restarts because it resides in a database outside the model.
- Semantic or exact lookup: Records are located through embeddings and similarity search, or through keys and keywords.
- Selective recall: Only a few highly relevant records are retrieved, which keeps the prompt compact and focused.
- Explicit management: Deciding what to keep, summarise or discard is part of context engineering.
How Memory in AI Agents Works
- Record each turn: Append the user message, the model response and every tool calling result to a message buffer.
- Enforce the limit: When the buffer exceeds its capacity, remove the oldest messages from the active window.
- Summarise removed turns: Condense the removed messages into a short summary that remains in the prompt.
- Persist important facts: Write durable outcomes, such as a confirmed root cause, to long-term storage with searchable keys or embeddings.
- Retrieve before reasoning: At the start of each task, query long-term memory with the current problem description.
- Assemble the prompt: Combine the system prompt, the retrieved records, the summary and the recent messages.
Example: Short-Term and Long-Term AI Agent Memory in Python
The program below maintains a bounded message buffer with a simple summary and recalls past incidents from a keyword-indexed store.
from collections import deque
class ShortTermMemory:
"""Keeps the last N messages; older ones are folded into a summary."""
def __init__(self, max_messages):
self.messages = deque(maxlen=max_messages)
self.summary = []
def add(self, role, text):
if len(self.messages) == self.messages.maxlen:
old_role, old_text = self.messages[0]
gist = " ".join(old_text.split()[:4]) # stand-in for an LLM summary
self.summary.append(f"{old_role}: {gist}")
self.messages.append((role, text))
class LongTermMemory:
"""Stores past incidents and finds them by shared keywords."""
def __init__(self):
self.records, self.index = [], {}
def save(self, text):
self.records.append(text)
for word in set(text.lower().split()):
self.index.setdefault(word, set()).add(len(self.records) - 1)
def search(self, query, top_k=1):
scores = {}
for word in set(query.lower().split()):
for i in self.index.get(word, ()):
scores[i] = scores.get(i, 0) + 1
ranked = sorted(scores.items(), key=lambda kv: (-kv[1], kv[0]))
return [(self.records[i], s) for i, s in ranked[:top_k]]
long_term = LongTermMemory()
long_term.save("INC-101 checkout-api timeout fixed by raising db pool size")
long_term.save("INC-117 search-api oom fixed by lowering batch size")
long_term.save("INC-130 checkout-api 5xx after deploy fixed by rollback")
short_term = ShortTermMemory(max_messages=3)
for role, text in [
("user", "Checkout alerts are firing again"),
("tool", "read_logs: checkout-api timeout on db connection"),
("agent", "Error rate is 4.2 percent over 15 minutes"),
("user", "Has this happened before?"),
]:
short_term.add(role, text)
print("Window:", [r for r, _ in short_term.messages])
print("Summary of dropped turns:", short_term.summary)
for record, score in long_term.search("checkout-api timeout db", top_k=2):
print(f"Recalled ({score} keyword matches): {record}")Window: ['tool', 'agent', 'user']
Summary of dropped turns: ['user: Checkout alerts are firing']
Recalled (3 keyword matches): INC-101 checkout-api timeout fixed by raising db pool size
Recalled (1 keyword matches): INC-130 checkout-api 5xx after deploy fixed by rollback- Bounded window: The buffer retains three messages, so the first user message is removed and survives only as a summary.
- Ranked recall: INC-101 shares three keywords with the query and is ranked above INC-130, which shares only one.
- Replaceable components: A production agent would substitute an LLM summariser and a vector database for these simplified versions.
Applications of Memory in AI Agents
- Incident response: Recalling previous root causes and remediation steps for recurring production alerts.
- Coding assistants: Remembering repository conventions, build commands and earlier review feedback.
- Long investigations: Preserving intermediate findings across many iterations of a ReAct agent loop.
- Task continuity: Resuming a multi-step plan after an interruption, as described in planning in AI agents.
- Personalisation: Storing each user's preferred output formats and notification channels.
Advantages
- Continuity: The agent continues a task without requesting the same information repeatedly.
- Accumulated experience: Resolved incidents become reusable knowledge for similar future problems.
- Token efficiency: Summaries and selective retrieval consume far fewer tokens than the complete history.
- Auditability: Stored records reveal which past information influenced a particular decision.
Limitations
- Information loss: Summarisation can discard a detail that later proves important.
- Irrelevant recall: Retrieval can return similar but unrelated records that mislead the model.
- Stale knowledge: Stored facts become outdated when systems, configurations or owners change.
- Privacy obligations: Persisted conversations may contain credentials or personal data that require retention controls.
- Memory poisoning: Malicious content written to long-term memory can influence many later sessions.
Quick Quiz
Pick an answer to check yourself. Nothing is saved.
Question 1 / 3
1. Where does short-term memory of an AI agent live?
Frequently Asked Questions
Do LLMs have memory?
Not between requests. A language model only sees the tokens sent in the current request, so any memory is provided by the application that stores and resends earlier information.
What is the difference between short-term and long-term memory in AI agents?
Short-term memory is the recent conversation kept inside the context window for the current task. Long-term memory is stored outside the model in a database and is retrieved only when it is relevant.
Is long-term memory in an AI agent the same as RAG?
They use the same retrieval technique, but the content differs. RAG usually searches a fixed document collection, while agent memory stores facts the agent itself has observed or been told.
How does an agent decide what to remember?
A write policy selects durable facts, such as a confirmed root cause or a user preference, and ignores routine messages. Some systems ask the model to extract these facts at the end of each session.
Related Articles
- Context Window in LLMLearn what the context window of an LLM is, how tokens fill it, what happens past the limit and how to fit long logs, with a Python truncation example.
- What is Context EngineeringLearn what context engineering is: selecting, ranking and fitting information into an LLM context window, with a Python token budget example and limits.
- Vector Database in RAGVector database in RAG explained: how embeddings are stored, indexed and searched with cosine similarity, ANN indexes, metadata filters and a Python demo.
- Tool Calling (Function Calling) in LLMLearn how tool calling works in LLMs: JSON Schema tool definitions, model tool calls, argument validation and tool results, with a Python DevOps example.
- ReAct Agent ExplainedLearn how a ReAct agent alternates thought, action and observation, with a Python example that traces a failing unit test to a commit and opens a ticket.
- Planning and Reasoning in AI AgentsLearn planning and reasoning in AI agents: task decomposition, dependencies, plan-and-execute and replanning, with a Python flaky test fix example.