AI AgentsAI Engineering

AI Agent Memory Explained: How Agents Remember (and Why Yours Forgets)

A practical map of AI agent memory: the four memory types, the wiring that makes each real (notes files, vector stores, memory layers), the seven reasons agents still forget, and a minimal setup you can copy today. Covers LangGraph, Letta, and Mem0.

Toolbit AI - Team
15 min read
AI Agent Memory Explained: How Agents Remember (and Why Yours Forgets)

Your agent was brilliant on Tuesday. It knew your project's naming conventions, remembered that you hate verbose logging, asked the right questions, and shipped the refactor. Then on Wednesday you opened a fresh chat, typed a quick follow-up, and got a polite, confused stranger.

If that stings, here is the comforting part: nothing is broken. The API your agent runs on is stateless. Every single request stands completely alone. The model has no idea Tuesday happened, because from its point of view, every conversation is its first conversation. Forgetting is not a bug you introduced. It is the default state of the technology, and memory is the feature you build on top of it.

And here is the even better part: agent memory is not one giant unsolved problem. It is four small, well-understood ones:

  • Working memory is the context window, everything the model can see during the current turn.
  • Episodic memory is what happened: past conversations, actions, and decisions.
  • Semantic memory is what it learned: distilled facts stored outside the model.
  • Procedural memory is how it behaves: rules, instructions, the system prompt and your code.

Out of the box, an agent has only the first one, which is exactly why it forgets. This post walks through all four: first the map, then the wiring for each type, then the real tools, then the seven reasons agents still forget even with memory installed, and finally a minimal setup you can copy this afternoon. No framework loyalty required.

In short:

  • The context window is working memory, not memory. It resets every turn unless you refill it.
  • Four types cover everything: working (the context window), episodic (what happened), semantic (what it learned), procedural (how it behaves).
  • Files are the cheapest long-term memory: a notes file the agent reads and writes beats most fancy setups.
  • Vector stores and memory layers earn their keep once facts outgrow a file.
  • Agents still forget for seven fixable reasons, from context rot to missing write paths.
  • In a hurry? Jump straight to the minimal memory setup near the end.

Agent memory, mapped: the four types

The four types of agent memory mapped to human memory concepts

Human memory research gives AI agents a working taxonomy: four agent memory types - working, episodic, semantic, and procedural - that the major frameworks all implement under different names.

When agent frameworks describe memory, they borrow the vocabulary of human memory, and they do it explicitly. LangChain's documentation states it directly: humans use memories to remember facts (semantic memory), experiences (episodic memory), and rules (procedural memory), and AI agents can use memory in the same ways. Anthropic supplies the fourth member of the family, calling the context window the model's "working memory": the text a language model can reference while generating a response.

Put the two together and you get a clean four-type map of agent memory:

  • Working memory: what is in front of the model right now. Exists only for the current turn.
  • Episodic memory: records of past events, actions, and decisions. The agent's diary.
  • Semantic memory: facts and knowledge distilled from those events. The agent's notebook.
  • Procedural memory: the rules that shape behavior. LangChain's docs describe it as the combination of model weights, agent code, and the prompt that together determine how the agent works.

Engineers then add a second naming layer on top, and this is where people get lost. LangGraph talks about short-term versus long-term memory. Letta talks about memory blocks, archival memory, and conversation search. These are the same four drawers with different labels on the front:

Cognitive termLangGraph calls itLetta calls itWhat it holds
WorkingShort-term memory: thread state plus a checkpointerMemory blocks and system/ files, pinned in contextThe conversation in flight right now
EpisodicThread history, resumable across sessionsConversation search over past chatsWhat happened: past turns, actions, decisions
SemanticLong-term store: JSON documents under a namespace and keyArchival memory, a semantic vector databaseLearned facts, retrieved by meaning
ProceduralPrompts, code, and model weightsThe system prompt and pinned core filesRules and behavior that guide every turn

Both vocabularies describe the same reality: one fast, tiny, expensive layer that exists right now, and three slower layers you own that persist between sessions. You can see the full taxonomy, with human examples for each type, in the LangChain memory overview, which is also where the episodic-versus-semantic framing enters most builder documentation (via the CoALA framework that the docs cite).

Working memory: the context window

How episodic and semantic memory wire into the context window

Working memory has a precise definition: every token the model can reference this turn. That means the system prompt, every message so far, all tool results, images, and documents, the definitions of every tool, and the new output and reasoning the model generates. All of it counts toward the window. There is no free space hiding anywhere.

The uncomfortable truth underneath is that the API is stateless. Each text generation request is independent, so multi-turn conversations only work because you provide the history again as parameters, every single turn. OpenAI's conversation state guide lays out the three options: replay the messages yourself, chain responses with previous_response_id, or use the Conversations API, which persists conversation state server-side as a long-running object with its own durable identifier.

So how big is working memory? Context windows now reach up to 1M tokens on some models. Surely a window that size fixes forgetting? No, and the reason has a name: context rot. As the number of tokens in the context window increases, the model's ability to accurately recall information from that context decreases, and this pattern shows up across all models. Context is a finite resource with diminishing returns: like humans, who have limited working memory capacity, models have an attention budget, and every token you add spends some of it.

Working memory is also the only memory you rent instead of own, because it is re-billed every turn. On GPT-5.2, input costs $1.75 per 1M tokens ($0.175 per 1M when cached, $14 per 1M for output). Some quick arithmetic, derived rather than a vendor claim: once your conversation has grown to 100k input tokens, every additional turn costs roughly $0.175 in input alone, before the model has done anything new. Prompt caching softens that bill, but cached tokens still occupy the window: caching changes what you pay for those tokens, not whether they count.

Vendors now ship features to stretch working memory: compaction (the server summarizes older turns), context editing (clearing stale tool results and old reasoning), and context awareness (the model tracks its own remaining token budget). These genuinely help, but notice what they are: ways to manage working memory. They do not replace the other three types. If you want to go deeper on treating the window as a budget, the context engineering skill that is quietly replacing prompt engineering covers the attention-budget idea in detail.

Episodic memory: files and transcripts

Episodic memory holds what actually happened: past conversations, actions taken, decisions made, and their outcomes. It is the agent's diary, and it is the first type of long-term memory most builders genuinely need.

Letta makes this an explicit product channel: its conversation search is historical retrieval, a way for the agent to recall what was said in previous conversations. Episodic memory, with a search box attached.

In practice there are two wiring patterns.

Pattern one: transcripts plus a checkpointer. LangGraph's short-term memory keeps message history inside a thread (a session-scoped conversation), then persists it to a database using a checkpointer, so the thread can be resumed at any time. You get pause-and-resume conversations, and the transcript itself is the memory: what happened, in order, replayable.

Pattern two: notes files. Anthropic describes a technique it calls structured note-taking, or agentic memory: the agent regularly writes notes that persist outside the context window, then pulls those notes back into the window at later times. The simplest version is an agent that maintains a NOTES.md file: decisions at the top, open questions below, updated as it works.

Letta has productized exactly this idea. In Letta, an agent's memory is its own git repository, projected onto disk through a memory filesystem called MemFS. Edits become memory once they are committed and pushed. Files under system/ sit in the system prompt every turn; everything else stays out of context, and the agent simply sees the file tree and reads what it needs. The agent's memory is literally files, versioned like code. You can read the full design in Letta's memory documentation.

Semantic memory: vector stores and memory layers

Episodic memory is a diary; semantic memory is an encyclopedia. It holds learned facts: things that are true about you, your project, your codebase, or the world, written intentionally rather than replayed verbatim. Where an episodic transcript says "the user asked me to use tabs at 14:02 last Thursday," semantic memory says "the user prefers tabs." Distilled, portable, and useful in every future session.

Two wiring patterns dominate here too.

Pattern one: vector stores. Letta's archival memory is a semantically searchable database where agents store facts, knowledge, and information for long-term retrieval. It cannot be pinned to the context window; the agent queries it on demand with dedicated tools, archival_memory_insert to write and archival_memory_search to read. Storage is unlimited, and it is agent-curated: only what the agent deliberately inserted is in there.

LangGraph takes the same shape with different words. Long-term memories are JSON documents in a store, each organized under a custom namespace (like a folder) and a key (like a file name), with an optional embedding index so you can search them by meaning rather than exact text. A profile under (user_id, "memories") holds that user's learned facts; a search across namespaces recalls them on demand.

Pattern two: memory layers, or extract-and-retrieve services. Instead of hand-wiring your own store, you run a dedicated memory system that your agent reads and writes through. Mem0's open-source offering is exactly this: a self-hosted memory management solution for AI agents and assistants, with Python and Node.js SDKs, that you host on your own servers with full control and no vendor lock-in.

There is also a timing question that trips people up: when should memories be written? LangGraph's documentation names the two options: in the hot path, meaning during agent logic as the conversation happens, or in the background, consolidated later. Letta gives the background option a charming name: dreaming, where the agent reflects on recent conversations and distills useful lessons into memory when the work is done.

One warning from the LangChain docs before you build: memory has shapes. A profile is a single, continuously updated JSON document about a user or organization. A collection is a growing set of separate memories. The docs caution that the profile approach can become error-prone as the profile gets larger, so plan for collections once memories multiply.

If the retrieval part of this is what you actually care about, you are closer to RAG than you might think: retrieval memory and RAG share the same machinery, so any solid guide to retrieval-augmented generation doubles as a guide to this wiring.

The memory tooling landscape: Letta, LangGraph, OpenAI, Mem0

Four real options, described by what each actually does, not by ranking:

ToolWhat it offers, per its own docsThe memory problem it solves
LettaAgent-owned git repo of memory files (MemFS), memory blocks and system/ files pinned in context, archival vector memory, conversation search"I want the agent's memory to be real files it owns, versioned, shareable via git remote on self-host"
LangChain / LangGraphShort-term memory via thread plus checkpointer; long-term store of JSON documents under namespaces; LangMem write and search tools"I want memory wired into my graph, from resumable threads to namespaced long-term facts"
OpenAIManual history replay, previous_response_id chaining, Conversations API durable state across sessions and devices, compaction endpoints"I want working-memory management handled at the API layer with minimal custom infrastructure"
Mem0Open-source self-hosted memory layer with Python and Node.js SDKs; managed platform from a free Hobby tier (10,000 add requests a month) through $19 Starter to $249 Pro per the official pricing page"I want memory as a service without building my own extraction and retrieval stack"

Pricing and plan details are as published by the vendor around September 2026 and can change - confirm on the official site.

The honest guidance line: pick by the problem you have. If you want resumable conversations, any checkpointer-style tool does it. If you want the agent to own a persistent, file-shaped identity, that is Letta's home turf. If you want memory inside your existing LangGraph graph, use the store. If you want to outsource the whole problem, that is the memory-layer pitch.

Why agents still forget

Here is the part nobody warns you about: wiring in memory does not automatically stop the forgetting. Seven distinct causes, each grounded in how the tools actually behave:

  1. Context rot. Dumping your entire memory store back into the context window recreates the original problem: as tokens grow, recall degrades, across all models. More memory in context is not better memory.
  2. Context pollution. Context windows of all sizes are subject to irrelevant information crowding out the relevant. Stale tool results and dead-end reasoning occupy the same budget as signal.
  3. Compaction loss. Overly aggressive summarization can discard subtle but critical context whose importance only becomes apparent later. The summary is not the transcript.
  4. No write path, or the wrong layer. Archival memory only contains what the agent deliberately inserted. If your agent never writes, the store is an empty library. And information that must always be visible belongs in pinned memory blocks, not in a store behind a search call.
  5. Statelessness at the API level. Each request is independent. If your harness does not resend history or persist state somewhere, the model starts from zero, no matter what memory system exists in your repository.
  6. Profile bloat. One ever-growing profile document becomes error-prone as it gets larger, per the LangChain docs. Memory needs shaping, not just storage.
  7. Retrieval is semantic, not exact. Vector search understands concepts and meaning rather than exact keywords, which is its superpower and its failure mode: a near-miss query can fail to surface the memory you know is in there.

The unifying lesson: memory that exists but never reaches the working memory this turn is functionally forgotten. A fact in a database the agent never queries might as well not exist. Anthropic's own documentation covers the rot and pollution failure modes in detail, and it is worth reading directly: context windows, by Anthropic.

A minimal memory setup you can copy today

You do not need a platform to get real memory. Here is a five-rung ladder, copyable in plain Python with any model, each rung useful on its own:

  1. Treat the context window as a budget. Assemble the smallest set of high-signal tokens that gets the job done: the current task, the essential history, the pinned rules. When you near the limit, compact. The major vendors now offer a way.
  2. Add a notes file. The agent writes structured notes, a NOTES.md-style file, outside the window, and reads them back at the start of each session. Episodic and semantic memory in one file, for the cost of a file. This is the cheapest real memory that exists.
  3. Add a retrieval store for facts. When notes outgrow one file, graduate to a store: LangGraph's store.put and store.search with namespaces, or Letta's archival tools. Same idea, indexed and queryable.
  4. Add a write policy. Decide explicitly: does the agent update memory in the hot path, mid-conversation, or in the background, after the work is done? Both are legitimate; the failure is never deciding.
  5. Upgrade path: self-host a memory layer. If building your own extraction is not the weekend you want, Mem0's open-source edition runs self-hosted with the agent writing and reading through its SDK.

That is the whole architecture: a budgeted window, a file for what happened and what was learned, a store for what must be findable, a policy for when to write. Everything else is polish. If you are wiring up your very first agent around this memory plan, start with any step-by-step agent-building walkthrough and bolt this memory plan onto it.

FAQ

Is the context window the same thing as agent memory? No. The context window is working memory only: it exists for the current turn and must be refilled every request. Long-term memory is separate infrastructure you build: files, transcripts, stores, and prompts that persist between sessions.

What is context rot, and why doesn't a 1M-token context window fix forgetting? Context rot is the finding that as token count in the window grows, the model's ability to accurately recall information from it decreases, and this shows up across all models. A bigger window holds more, but recall still degrades as it fills, so a million tokens do not buy you a perfect memory.

How much does it cost to keep resending conversation history every turn? Working memory is re-billed every turn. On GPT-5.2 at $1.75 per 1M input tokens, a conversation grown to 100k input tokens costs roughly $0.175 per turn in input alone (derived arithmetic). Prompt caching cuts the rate to a tenth, but cached tokens still occupy the window.

What is the cheapest memory system that actually works? A markdown notes file the agent reads and writes: structured note-taking that persists outside the context window and gets pulled back in later. Letta's MemFS productizes the same idea as a git-backed memory filesystem.

Why does my agent forget things even with RAG wired in? Retrieval is only half a memory system. If the agent has no write path, the store is empty; if queries near-miss, semantic search fails to surface what is in there; if you stuff everything into context, rot sets in; and compaction can silently drop the critical detail. Memory must both exist and reach the working memory this turn.

Share this article

Related articles

Continue exploring similar guides and insights