AI agent memory is whatever your code loads back into the prompt on each call. Start with a window of recent turns, add a rolling summary for long sessions, a small per-user profile for facts that must survive between sessions, and vector recall only when the memory is large and unstructured. On published token prices the options sit within about 30 percent of each other; replaying the full history is the expensive one unless prompt caching applies.
AI agent memory is whatever your code stores between model calls and loads back into the prompt, because the model itself keeps nothing from one request to the next. There are four practical options: a window of recent turns, a rolling summary, vector recall and an entity profile. Most agents need the window plus one of the others, and plenty never need a vector database.
Checked October 2026 against the LangChain memory overview and short-term memory guide, OpenAI's conversation state and prompt caching guides and pricing page, Anthropic's memory tool and compaction docs, the n8n memory node docs, the pgvector README, and the Mem0 and Zep pricing pages.
This is a documentation-based comparison. The costs are calculated from published prices with inputs we state in full, so you can swap in your own; we did not benchmark recall quality. It sits under our agentic AI workflows hub, which covers the frameworks and guardrails around the memory layer.
Which AI agent memory should you pick?
Pick by two questions: does the agent need to remember within one conversation or across many, and can you look the memory up exactly or do you have to compress or search it?
Start top right. Add another cell only when a real conversation fails without it.
- Window only if conversations are short and self-contained: support chat, booking, form filling.
- Add a summary if single sessions run long (research, coding, multi-step tasks) and early turns still matter at the end.
- Add an entity profile if the agent should know a returning user: name, plan, preferences, the order they asked about last week.
- Add vector recall when what must be remembered is large, unstructured and has no key to look it up by: months of tickets, call notes, documents.
Agent memory types: window, summary, vector and entity side by side
Each type is a different answer to the same question: which old tokens earn a place in the next prompt.
| Window | Summary | Vector | Entity | |
|---|---|---|---|---|
| What goes in the prompt | The last N turns, word for word | A running summary plus the newest turns | The few stored snippets most similar to the current message | A short profile of facts about the user or task |
| What it forgets | Everything older than N turns | Details the summariser dropped | Anything the search does not surface | Anything outside the fields you planned |
| Extra work | None | A summariser call when history passes a threshold | An embedding call and a similarity query per turn | An extraction call, per turn or in the background |
| Stored as | Rows keyed by thread ID | One summary per thread | Vectors in pgvector or a vector database | One JSON document per user |
| Scope | One conversation | One conversation | Across conversations | Across conversations |
| Breaks when | The key fact was said 20 turns ago | Exact wording, numbers or IDs matter | The answer needs a lookup by key, order or date | Users say things you had no field for |
The four names come from LangChain's original memory classes: ConversationBufferWindowMemory, ConversationSummaryMemory, VectorStoreRetrieverMemory and ConversationEntityMemory. In the current LangChain repository all four sit in the langchain_classic package and are marked deprecated, with a pointer to checkpointers and stores instead. The ideas outlived the classes.
LangChain's current docs sort memory by scope. Short-term memory belongs to one thread and is saved by a checkpointer, so window and summary live there. Long-term memory lives in a store under a namespace you choose and can be read from any thread. The same docs borrow three labels from psychology: semantic memory for facts, episodic for past experiences and procedural for instructions. An entity profile is semantic memory kept as one document.
How does LLM memory management work on each call?
Every turn runs the same loop, whatever framework you use. Memory types differ only in the load and write steps.
- 01Load
Fetch recent turns by thread ID, the profile by user ID, and any retrieved snippets.
- 02Assemble
System prompt, then memory, then the new message. Order matters for caching.
- 03Call the model
You pay input tokens for everything assembled, every turn.
- 04Write
Append the new turn. Update the profile or embed the turn if you use those.
- 05Compact
When history passes a threshold, trim it or fold the oldest turns into a summary.
Providers now handle parts of this loop for you, which saves code but not tokens:
- OpenAI. The Responses API chains turns with
previous_response_idor a durable conversation object. Response objects are kept for 30 days by default; items attached to a conversation are not subject to that limit. The docs are explicit that all previous input tokens in the chain are still billed as input. Server-side compaction can shrink the context once it crosses acompact_threshold. - Anthropic. Compaction (beta) replaces older turns with a summary Claude writes on the server. The memory tool is a file-based long-term memory: Claude requests file operations under
/memoriesand your application executes them against storage you control. - LangChain and LangGraph. A checkpointer (
InMemorySaverfor development,PostgresSaverin production) persists the thread;SummarizationMiddlewaresummarises once a token trigger is hit and keeps the most recent messages. - n8n. Memory sub-nodes plug into the AI Agent node. All of them are windows with a Context Window Length. Our n8n and OpenAI guide shows the agent node they attach to.
What each memory type costs per 1,000 conversations
Memory cost is almost entirely input tokens: whatever you load is billed again on every turn. The table prices one conversation shape on two OpenAI models. This is arithmetic from published rates, not a measurement, and every input is an assumption to replace with your own.
| Per 1,000 conversations | 20 turns, GPT-6 Luna | 20 turns, GPT-6.1 Sol | 100 turns, GPT-6 Luna | 100 turns, GPT-6.1 Sol |
|---|---|---|---|---|
| Full history, list price | $9.26 | $185.20 | $166.30 | $3,326.00 |
| Window (last 5 turns) | $6.11 | $122.20 | $32.35 | $647.00 |
| Window + rolling summary | $7.97 | $159.40 | $46.93 | $938.60 |
| Window + vector recall | $7.46 | $146.35 | $39.11 | $767.76 |
| Window + entity profile | $7.25 | $144.90 | $37.48 | $749.70 |
| Full history, best-case caching | $3.79 | $69.53 | $30.17 | $451.37 |
Source: Calculated from the OpenAI pricing page, checked October 2026. Illustrative inputs, not measured.
Four things stand out.
- The add-ons are cheap. A summary, vector recall or a profile adds 19 to 30 percent to a plain window at 20 turns. On a small model that is one or two dollars per 1,000 conversations. Choose between them on what the agent needs to remember, not on price.
- Embeddings are a rounding error. Embedding every message and query in a 20-turn conversation is 7,600 tokens, about 15 cents per 1,000 conversations. The cost of vector recall is the snippets you inject, not the vectors.
- Full history grows with the square of the length. At 20 turns it costs about 50 percent more than the window. At 100 turns it costs about five times as much, and the model is reading 30,000 tokens of chat to answer one line.
- The model matters more than the memory. Every list-price row is 20 times higher on GPT-6.1 Sol than on GPT-6 Luna. Our LLM API pricing comparison lists the current rates across providers.
Prompt caching changes the ranking
The last row of the table is the surprise: replaying the full history can be the cheapest option while a conversation is short, if the prompt cache hits. OpenAI's prompt caching guide (checked October 2026) says that on GPT-5.6 and later models a prefix of at least 1,024 tokens is written to the cache at 1.25 times the input rate and read back at 0.1 times, or 0.05 times on GPT-6.1 Sol, and stays eligible for 30 minutes after its last use.
Caching only works on an unchanged prefix. A full history is append-only, so each turn reads everything before it at the cached rate. A sliding window drops its oldest turn on every call, so the text after the system prompt never matches and nothing is reused. That is why the best-case cached history comes out at $69.53 against $122.20 for the window in the 20-turn example. It is a best case: it assumes turns arrive within the cache lifetime and reach a machine that holds the entry, and OpenAI is clear that a hit is never certain.
Window memory: the default
A window replays the last N turns and drops the rest. It needs no extra calls and is predictable to bill. In n8n this is every memory node: Simple Memory, Postgres Chat Memory, Redis Chat Memory and MongoDB Chat Memory all expose a Context Window Length. The n8n docs warn that Simple Memory does not work in queue mode, because calls may land on different workers, so use the Postgres or Redis node in production (our n8n Postgres guide covers the credentials). The Zep and Motorhead memory nodes are marked deprecated.
The limit is the cliff. Turn N+1 is gone, however important it was. If users give an order number at the start and ask about it ten minutes later, a five-turn window will ask them for it again.
Summary memory: for long single sessions
Summary memory folds older turns into a running note and keeps the newest turns verbatim. It removes the cliff at the price of an extra model call each time it compacts, and it carries more tokens per turn than a bare window, which is why it sits above the window in the cost table. You can run the summariser on a cheaper model than the agent.
Summaries lose exact values. Names, amounts, IDs and dates are the first things a summariser paraphrases away, so pair it with an entity profile when those matter. Both OpenAI and Anthropic now offer server-side compaction, so you may not need to write the summariser yourself.
Vector memory: recall by meaning
Vector memory embeds past content, stores the vectors and retrieves the closest matches to the current message. It is the only option here that scales to history far larger than the context window, and the only one that finds something phrased differently from how it was stored.
It also fails in ways the others do not. Similarity search cannot answer "what did I ask first" or "what changed since Tuesday", it returns near-duplicates unless you deduplicate on write, and a stale memory looks as relevant as a current one. In n8n, connect the PGVector Vector Store node to the agent as a tool rather than looking for a vector memory node.
Entity memory: a profile, not a transcript
Entity memory keeps a short structured record per user or task and updates it as facts appear. LangChain's docs describe two shapes: a single profile document you rewrite, or a collection of small memories you add to. Profiles are easy to load and easy to show the user; collections lose less but need search and cleanup.
You also choose when to write. On the hot path the agent saves a fact before it replies, which adds latency. In the background a separate job reads the finished conversation and updates the profile, which is what the cost table assumes. Anthropic's memory tool is a file-based take on the same idea, and a natural fit if you already use Claude tool calling.
When is a vector database overkill for agent memory?
More often than tutorials suggest. The recurring complaint in agent-building communities, that agents do not always need a vector database, holds up against the numbers: retrieval is not expensive, but it is a second system to run, tune and debug, and most memory problems are lookups, not searches.
- Everything worth remembering about a user fits in a few hundred tokens
- The memory has a key: user ID, order number, date, ticket ID
- Conversations are short enough to replay or summarise
- You have a few thousand rows, where exact search in pgvector needs no index
- Nobody has shown you a conversation that failed for lack of recall
- History per user is far larger than the context window
- The content is unstructured: notes, tickets, transcripts, documents
- Users ask about old things in new words
- You cannot predict which past detail will matter
- You can filter by user or date first, then search within that slice
When you do need it, start inside Postgres. The pgvector README notes that it performs exact nearest-neighbour search by default, with perfect recall, and that HNSW and IVFFlat indexes trade some recall for speed once tables get large. That is enough for most agents; our Supabase tutorial covers setting up the Postgres side.
Which AI agent memory database should you use?
One Postgres database covers all four memory types. Reach for something else when you have a specific reason. Prices below were checked in October 2026.
| Option | Best for | What the docs say | Trade-off |
|---|---|---|---|
| Postgres | Window, summary and profile in plain tables | Already in most stacks; n8n and LangGraph both ship Postgres memory | You write the trimming and extraction logic |
| Postgres + pgvector | Vector recall next to the rest | Exact search by default; HNSW and IVFFlat indexes when tables grow | Indexes cover vectors up to 2,000 dimensions |
| Redis | Short sessions that should expire | n8n Redis Chat Memory has a session time-to-live setting | Not where you keep long-term facts |
| SQLite | Local agents and prototypes | The quick-start session store in the OpenAI Agents SDK | Single machine |
| Mem0 | Managed fact extraction and retrieval | Free tier; Starter $19/mo; Pro $249/mo adds graph memory | Metered on add and retrieval requests |
| Zep | Managed context graph with temporal memory | Free 10,000 credits/mo; Flex $125/mo with 50,000 credits | Metered on what you write; retrieval is unmetered |
Managed memory services bill by request, not by token, so run the same arithmetic. With one write and one retrieval per turn, 1,000 conversations of 20 turns is 20,000 retrieval requests a month: more than Mem0's Starter plan includes (5,000) and inside Pro (50,000). On Zep, where an episode costs one credit per 350 bytes, sending each message as an episode is roughly 80,000 credits at about four characters per token: the $125 Flex plan plus about $75 of overage. Both are estimates from the published plan limits.
Our default stack
This is the order we would add memory to a new agent, and the point at which each layer becomes worth its complexity. It is a recommendation, not a test result.
- Day oneWindow in Postgres
Last 5 to 10 turns by thread ID. Stable system prompt first so it can be cached.
- Returning usersEntity profile
One JSON row per user, updated by a background job after each conversation.
- Long sessionsThreshold compaction
Summarise when history passes a token limit. Use the provider feature if it exists.
- Proven needpgvector in the same database
Only after real conversations fail on recall. Filter by user before you search.
- At scaleA dedicated memory service
When extraction, deduplication and time-aware facts have become a project of their own.
The first, second and fourth layers fit in three tables. Frameworks create their own versions (n8n's Postgres Chat Memory creates its table if it does not exist, and LangGraph's checkpointer has a setup() call), so treat this as the shape, not a migration to copy.
-- 1. Window: recent turns, loaded by thread
create table chat_messages (
id bigserial primary key,
thread_id text not null,
role text not null,
content text not null,
created_at timestamptz default now()
);
create index on chat_messages (thread_id, created_at);
-- 2. Entity memory: one small profile per user
create table user_profiles (
user_id text primary key,
profile jsonb not null default '{}',
updated_at timestamptz default now()
);
-- 3. Vector recall: add only when you need it
create extension vector;
create table memories (
id bigserial primary key,
user_id text not null,
content text not null,
embedding vector(1536)
);
-- the four nearest memories for one user, by cosine distance
select content from memories
where user_id = $1
order by embedding <=> $2
limit 4;The vector(1536) column matches the default size of OpenAI's text-embedding-3-small; change it to match your embedding model. If you are building memory into a product rather than a one-off agent, our AI SaaS Builder program goes further on the parts this page only prices: prompt caching and cost control on the Claude API, and retrieval with embeddings and pgvector on Supabase.
For where memory sits among tools, triggers and approvals, read the AI agent automation guide; if several agents share state, see agent-to-agent workflows in n8n.
AI agent memory: FAQ
What is AI agent memory?
AI agent memory is the data an application stores between model calls and loads back into the prompt, because a language model keeps nothing from one request to the next. In practice it is some mix of recent conversation turns, a summary of older ones, facts about the user and snippets retrieved by search. The agent only remembers what your code puts back in front of it.
What are the main agent memory types?
Four cover most builds. Window memory replays the last few turns. Summary memory replaces older turns with a running summary. Vector memory embeds past content and retrieves the snippets most similar to the current message. Entity memory keeps a small structured profile of facts about a user or task. LangChain groups the first two as short-term, thread-scoped memory and the last two as long-term memory.
Does an AI agent need a vector database for memory?
Often not. If conversations are short, a window is enough. If the agent must remember a user between sessions, a small profile loaded by user ID usually covers it. A vector store earns its place when the memory is large, unstructured and cannot be looked up by a key, such as months of tickets or documents. Even then, pgvector inside the Postgres you already run is a reasonable first step.
How much does AI agent memory cost?
Mostly input tokens. In our illustrative 20-turn conversation priced on OpenAI list rates (checked October 2026), a five-turn window cost about $6 per 1,000 conversations on GPT-6 Luna and about $122 on GPT-6.1 Sol. Adding a summary, vector recall or a profile added roughly 19 to 30 percent. Replaying the full history cost about 50 percent more than the window, before prompt caching.
What is the best database for AI agent memory?
Postgres is the practical default: one table holds chat history by thread, a JSON column holds user profiles, and the pgvector extension adds similarity search when you need it. Redis suits short-lived sessions that should expire. Dedicated memory services such as Mem0 or Zep make sense when you would otherwise build fact extraction, deduplication and graph relationships yourself.
Does the OpenAI API remember the conversation for me?
It can store it, but you still pay for it. The Responses API chains turns with previous_response_id or a conversation object, so you do not resend history yourself. OpenAI documents that all previous input tokens in the chain are still billed as input tokens. Server-side state saves you code and storage, and the token bill is the same as replaying the full history.
What memory options does n8n have for AI agents?
The n8n docs list Simple Memory, Postgres Chat Memory, Redis Chat Memory, MongoDB Chat Memory and Xata as memory sub-nodes, plus a Chat Memory Manager node. All of them are window memories with a Context Window Length setting. Simple Memory does not work in queue mode. For recall across sessions, connect a vector store node such as PGVector to the agent as a tool.
Memory is one layer. Build the product around it.
AI SaaS Builder, included in All Access, covers the Claude API, cost control with prompt caching, and retrieval with embeddings and pgvector on Supabase, alongside the other three programs, live coaching and the private community.
Compare memory setups with other builders
Join the free Discord to share what your agent remembers, what it forgets and what it costs you per conversation.