Appearance
Memory Systems
An LLM is stateless on its own: each call sees only the tokens in that one request, and the previous conversation might as well never have happened. When an agent "remembers you," 100% of that ability comes from the harness—it is the harness that decides, at every step, which pieces of history get pushed back into the context window. A memory system is the set of mechanisms inside the harness responsible for persisting, filtering, and injecting that history.
The one-line definition
Memory system = the full set of mechanisms that decide which information is worth surviving across the boundary of the context window, and in what form it comes back to context next time.
Why you need a memory system
On the surface the answer is obvious: the context window can't hold all the history. But the real engineering motivations stack up in three layers, each more critical than the last:
- Capacity constraints. Even with a 1M-token context window, filling it is bad: cost grows linearly, latency climbs, and the model's attention to the middle of a long context degrades (the "lost in the middle" phenomenon). Memory's first job is not "store more" but "make sure the model only sees what's worth seeing."
- Cross-session continuity. Software engineering tasks span days; user preferences span months. Without persistent memory, every session is a cold start—the user has to re-explain the project structure, build commands, and code style. This is one of the biggest sources of variance in the coding-agent experience.
- Accumulated experience. Much of a human engineer's value comes from "I fell into that pit last time." If an agent can bank lessons like "this repo's tests must run under xvfb," its capability compounds with every use; otherwise every mistake gets repeated. This is the most underrated function of a memory system.
Memory is not free
Every piece of memory injected into context eats into the working-memory budget and raises the risk of interference. The design goal of a memory system is never "remember everything"—it is to buy the maximum continuity benefit at the minimum context cost.
The three-layer memory model
Borrowing the layered thinking of cognitive science and operating systems, memory systems that work in practice almost all converge on a three-tier structure:
┌────────────────────────────────────────────────────────────────────────────┐
│ L1 In-context working memory │
│ Message history, tool results, and scratchpad for the current session │
│ Latency: 0 (already in context) Capacity: limited by context window │
│ Eviction: compression / truncation / summarization │
├────────────────────────────────────────────────────────────────────────────┤
│ L2 Session-level memory │
│ Session summaries, task state, todo lists, interim conclusions │
│ Latency: very low (file or small data) Capacity: KB-scale │
│ Eviction: session archiving; discarded once a task completes │
├────────────────────────────────────────────────────────────────────────────┤
│ L3 Cross-session long-term memory │
│ User preferences, project conventions, lessons learned, domain knowledge│
│ Latency: explicit load or retrieval Capacity: unbounded in theory │
│ Eviction: manual editing / conflict merging / periodic cleanup │
└────────────────────────────────────────────────────────────────────────────┘The core difference between the three tiers is not how long things are stored, but how they enter context:
| Layer | How it enters context | Typical carrier | Who writes |
|---|---|---|---|
| L1 Working memory | Naturally present | Message list | Appended automatically by the harness |
| L2 Session memory | Injected after compression | Summaries, todo files | Harness automatically + agent proactively |
| L3 Long-term memory | Explicitly loaded or retrieved | Memory files, databases, vector databases | User handwritten + agent learned |
Managing L1 (compression, truncation, summarization) falls under context engineering; this chapter focuses on L2 and L3—the parts that genuinely need to be "designed."
The memory-file pattern: from CLAUDE.md to AGENTS.md
From 2024 to 2026, long-term memory for coding agents converged on a remarkably uniform answer: plain Markdown files, kept in the repository, injected wholesale at session start.
What it looks like
Claude Code's memory files come in three scopes, all loaded and concatenated at session startup:
- Enterprise policy: managed by organization administrators, applies to every project in the organization, highest priority;
- Project memory:
./CLAUDE.mdat the repo root, committed to git, shared across the whole team—the home for build commands, code conventions, and architecture decisions; - User memory:
~/.claude/CLAUDE.md, personal preferences that apply across all projects.
It also supports nested CLAUDE.md files (subdirectory files loaded on demand), the @path/to/file.md import syntax (pull in other files' contents while keeping the main file tidy), the /memory command for editing memory files directly, and typing # at the start of input to quickly append a memory. The early CLAUDE.local.md (personal, project-level, not committed) has been deprecated; the official advice is to use the import syntax instead.
AGENTS.md turns the same idea into a cross-vendor open convention: a single Markdown file at the repo root, "a README for agents." According to its website, the format is championed jointly by OpenAI Codex, Amp, Google Jules, Cursor, Factory, and others, has been adopted by more than 60,000 open source projects, and is now hosted by the Agentic AI Foundation under the Linux Foundation. The rules are equally plain: the AGENTS.md closest to the file being edited wins, and explicit user instructions in the conversation override any file.
Why plain text files beat databases
This convergence is no accident—it reveals the actual structure of what long-term memory is asked to hold:
- Readable, editable, diffable. A memory file is ordinary Markdown: users can see directly what the agent "believes," track every change in git, and delete it in one step. Retrieval-based approaches (vector databases) fundamentally can't do this—there is no way to audit the set of "things the retriever thinks are relevant."
- The content is instructions and conventions, not mountains of facts. What a coding agent needs to remember is stuff like "this project uses pnpm, not npm" and "lint must pass before committing"—a small, high-value, highly stable set of conventions. At this scale (a few hundred lines), injecting the whole thing is cheap; retrieval would be over-engineering.
- Natural support for hierarchy and proximity. Nested files give each subproject in a monorepo its own conventions, and "closest file wins" is a zero-config disambiguation rule.
The decision criterion
When the content of long-term memory is conventions and preferences (few, stable, needing user audit), choose memory files. Only when the content is historical facts and experiences (many, messy, recalled on demand) do you need a retrieval-based approach. Most coding agents' primary need is the former.
The memory directory: memory the agent writes itself
The next evolutionary step for memory files is transferring part of the write authority from humans to the agent. Taking Claude Code as an example: besides the user-maintained CLAUDE.md, community sources indicate that v2.1.59 (February 2026) introduced auto memory, on by default: while working, the agent autonomously writes down information worth keeping (debugging findings, preferences the user has corrected, key decisions) into a memory directory isolated per project (by default ~/.claude/projects/<project-path>/memory/), which later sessions recall automatically and which can be viewed and cleaned up with /memory. This directory stays out of git and is private to the local machine, forming a two-layer structure alongside the team-shared CLAUDE.md: shared conventions written by humans plus private notes written by the agent.
This division of labor is worth remembering: shared, auditable content goes in in-repo files; private, stream-of-work content goes in the local directory. Both are plain text, and both preserve inspectability.
When to write: what's worth remembering
The hard part of a memory system is not storing—it's deciding what to write. Write too little and the agent stays a perpetual newcomer; write too much and the memory store becomes a graveyard of noise, polluting retrieval and injection with irrelevant content.
Write triggers that have proven effective in practice:
| Worth remembering | Not worth remembering |
|---|---|
| Explicit user corrections ("don't use X, use Y") | Things the model guessed right on the first try |
Non-obvious facts that took repeated attempts to solve ("this test needs --no-sandbox") | Facts a single read of the docs would surface |
| Stable preferences the user has voiced (code style, communication style) | One-off task details ("rename this variable to foo") |
| Implicit project-specific conventions (documented nowhere) | Anything already written in the README / CLAUDE.md |
| Environmental traps that caused failures | Chitchat unrelated to the current task |
The writing criterion boils down to one rule: only record information whose absence would cause a mistake or repeated work next time. A useful self-check: if this piece of information were deleted tomorrow, would it cause a foreseeable loss? If not, don't write it.
Who does the writing also matters:
- User-triggered (
#quick memory, "remember this"): the most reliable and the highest signal-to-noise ratio; always support it. - Agent-initiated writes: these need explicit writing criteria in the system prompt, otherwise the model either records nothing or pastes in tool output wholesale. Letta's approach is to give each memory block a
descriptionfield telling the agent what belongs in that block—the better the description, the sharper the agent's judgment about what to write. - Harness-automatic capture (e.g., generating a summary at session end): fits L2, but keep the granularity under control—summaries of summaries distort exponentially.
Forgetting and eviction: the badly underrated half
A memory system that only takes in and never lets go will inevitably rot. In real-world scenarios, forgetting strategies include at least four mechanisms:
- Conflict merging. When a new memory contradicts an old one ("I'm training for a marathon" vs. "I sprained my ankle"), either have the model actively rewrite the old entry at write time, or run a consistency check before injection. OpenAI explicitly noted in the new ChatGPT memory system that one problem with the old "saved memories" mechanism was entries contradicting each other; the new version switches to a continuously updated "memory summary" to mitigate this.
- Capacity limits. Each Letta memory block has a character cap (e.g., 5,000 characters); once full, the agent must itself decide what to compress or replace—turning eviction into an explicit action by the agent rather than silent dropping.
- Manual cleanup channels. Editing via
/memory, manual deletion on ChatGPT's memory summary page—these are necessary facilities. Users must be able to answer "what exactly does it remember about me?" - Scope isolation instead of deletion. Project-level memory must not leak across projects; temporary sessions (like ChatGPT's Temporary Chat) neither read nor write memory at all—many "forgetting" requirements are really isolation requirements.
Don't treat vector similarity as an eviction policy
Fully automatic eviction schemes like "decay over time" or "deduplicate by similarity" sound elegant, but in practice they delete memories that are infrequent yet critical (like a release process used once a year). Eviction decisions should go either to a human or to the model making an explicit judgment—not to statistical heuristics.
Memory vs. RAG: not the same thing
Memory systems and RAG (retrieval-augmented generation) are constantly conflated in engineering practice, because both may use vector retrieval. But their objective functions differ:
| Dimension | Memory system | RAG |
|---|---|---|
| Content source | The agent's own experiences and interaction artifacts | External corpora (docs, codebases, web pages) |
| Core question | Which interaction artifacts are worth banking, and when to inject them | How to recall relevant chunks from a large corpus |
| Who writes | User + agent (grows dynamically) | Usually bulk-loaded offline (relatively static) |
| Personalization | Strong, oriented to "this user / this project" | Weak, oriented to all users |
| Failure modes | Wrong information remembered, memory contamination | Inaccurate recall, poor chunking |
How the two fit together is equally clear: RAG handles "knowledge of the world"; memory handles "the history between you and me." A coding agent uses RAG to search codebases and documentation (see Tool Systems and Context Engineering) and a memory system to store project conventions and user preferences. Sharing one vector infrastructure is fine, but the write paths, eviction policies, and injection logic should be designed separately—mixing user preferences and code snippets in the same vector database is a classic breeding ground for "memory contamination" failures.
The pitfalls of vector-retrieval memory
Vector retrieval remains an important option for long-term memory (especially for conversation history and large-scale experience recall), but this direction has a set of recurring pitfalls:
- Similarity ≠ relevance. Embeddings measure semantic similarity, but memory needs "useful for the current task." When a user says "help me look at that bug," vector retrieval will recall a pile of memories that talk about bugs, but not necessarily the one about "how we fixed that bug last time." Temporal order, causality, and task context are things embeddings cannot express.
- No review at write time, unpredictable results at read time. Under the memory-file pattern, users can see all the memory; a vector database holding thousands of entries is a black box—users have no idea what's inside, and cannot predict what any given query will dredge up. The loss of auditability is the biggest hidden cost of the vector approach.
- Degradation and drift. As writes accumulate, semantic near-neighbors in the store get denser and the signal-to-noise ratio of top-k retrieval keeps falling. A vector memory store without an eviction mechanism often performs worse six months in than it did when first built.
- Hidden upgrade costs of embedding models. Switching embedding models means reindexing the entire store, and memory data is precisely the hardest kind to migrate—it is full of fragments that only make sense with their original context.
- "Retrieved but ignored." Even when recall is correct, a poor injection position (buried in the middle of tens of thousands of tokens) can get the memory overlooked by the model. Retrieval is only half of the memory pipeline; the injection strategy matters just as much.
This is not to say vector memory is unusable—Letta's archival memory is vector-based, and retrieval sits behind ChatGPT's memory summary too. The point is: vector retrieval is a good recall channel for "experiential" memory in L3; it is not suited to being the only memory mechanism, and even less suited to replacing auditable memory files.
Three representative implementations
Claude Code: files as memory
Claude Code represents the "minimal and auditable" route: long-term memory is just layered Markdown files (enterprise/project/user three-tier CLAUDE.md plus nested files plus @ imports), read in wholesale at session start, plus a locally maintained memory directory where the agent deposits private notes. The whole system has no vector index—an implementation analysis circulating in the community points out that its selection among memory files relies on an LLM reading file headers to make the call, not on embedding similarity. The design philosophy is explicit: memory that users can see, can edit, and can commit to git is the only memory you can trust. The cost is capacity—wholesale injection means memory files must stay lean; a few hundred lines is a healthy size, and a CLAUDE.md that runs to thousands of lines will dilute the critical instructions. For a deep dive, see Case Study: Claude Code.
ChatGPT memory: a continuously updated user profile
ChatGPT's memory takes the "automatic user profile" route: in the background, the system continuously synthesizes the facts revealed in conversation (occupation, preferences, project background) into a continuously updated "memory summary," replacing the early saved memories that users had to manage by hand. The improvements OpenAI's official docs emphasize happen to confirm two principles of this chapter: the old system's problem was that entries go stale and contradict each other (merging and eviction were missing); the new system's focus is visibility—users can see, edit, and delete memories on the summary page, and can click the "sources" icon below an answer to see which memories it drew on. Temporary sessions (Temporary Chat) provide a thorough isolation option. This is the most complete sample of "memory governance" in a consumer product.
Letta (formerly MemGPT): OS memory management for LLMs
Letta's predecessor was MemGPT, released in October 2023 (arXiv:2310.08560), whose core idea was "virtual context management": the way an OS pages memory between fast and slow storage, it moves information between a constrained context window and external storage, and uses an "interrupt" mechanism to manage control flow between the agent and the user.
In the Letta product, this idea materializes as a clear set of layers:
- Memory blocks: segments of context pinned in the system prompt as XML structure, always visible. Each block has four fields—
label,description,value, andlimit(a character cap)—and the agent reads and writes them autonomously through built-in memory tools; blocks can be made read-only (read_only), and can be mounted by multiple agents to implement shared memory. The classic configuration is a pair of blocks:persona(the agent's persona) andhuman(the user's profile). - Out-of-context storage: all messages and state are persisted in a database; nothing squeezed out of the context window is lost, and the agent can find it again through retrieval tools—corresponding to MemGPT's archival / recall memory concepts.
Letta's contribution is proving the viability of "agents managing their own memory": writing, compression, and eviction are all done by the model through tool calls, while the harness provides only the structure and the limits. The lesson it offers for memory-system design is universal—adding the three metadata fields label, description, and limit to memory is the minimum-cost upgrade from "unstructured notes" to "manageable memory."
Designing a memory system: a decision checklist
Before you build, answer these seven questions in order:
- What to remember? List the memory types your agent actually needs: user preferences, project conventions, task state, conversation history, domain facts. Different types call for different carriers.
- Which layer does each type belong to? Map them to L1/L2/L3. Content that can be injected wholesale and needs auditing goes into memory files; only content that is voluminous and recalled on demand justifies retrieval.
- Who gets write authority? User handwritten, agent initiated, and harness automatic can coexist, but each memory type must have one clearly designated writer; otherwise conflicts are unresolvable.
- What is the writing standard? If you let the agent write, give it a hard criterion in the prompt—like "only record information whose absence would cause a mistake next time"—and write good block descriptions.
- How does eviction work? What are the limits? Once full, who decides what gets deleted? Who merges conflicts between old and new?—if you can't answer, start with files plus manual editing.
- Can users see and delete it? Memory must be auditable and erasable. Any design that can't do this (black-box vector databases) needs a strong justification.
- How is it isolated? Where are the memory boundaries between projects, between users, and for temporary sessions? Isolation often matters more than deletion.
The minimum viable setup
A first-version memory system needs only two things: an in-repo AGENTS.md (conventions written by humans) plus a session-end summary file (written automatically by the harness). These two cover 80% of the value with zero infrastructure. Add retrieval only when a concrete problem (cross-session experience recall) backs you into a corner.
Further reading
- Context Engineering: managing L1 working memory—compression, truncation, and injection strategies
- Agent Loop: where in the loop memory reads and writes happen
- Skill System: the division of labor between skills and memory—reusable procedures vs. reusable facts
- Case Study: Claude Code: a complete anatomy of file-based memory
- Case Study: LangGraph: framework-managed state and checkpoint memory
- Practice: Common Pitfalls: anti-patterns like memory contamination and context bloat
- Papers: Core Papers: the lineage of memory papers, from MemGPT onward
References
- Claude Code Memory & CLAUDE.md guide (three scopes and import syntax)
- Claude Code Memory Explained: CLAUDE.md and Auto Memory (auto memory version info)
- Claude Code Memory: auto memory storage path notes
- AGENTS.md website (format spec, adoption scale, hosting)
- OpenAI official Memory FAQ (memory summary, saved memories, Temporary Chat)
- MemGPT: Towards LLMs as Operating Systems (arXiv:2310.08560)
- Letta official docs: stateful agents and memory concepts
- Letta official docs: memory blocks (label/description/limit/read_only)
Verification note
The official Claude Code memory documentation (code.claude.com/docs/en/memory) could not be fetched from the web at writing time; the statements above about CLAUDE.md scopes, the # shortcut, the /memory command, and auto memory were cross-checked against the secondary sources listed above. The specific auto memory version number (v2.1.59) appears only in a third-party blog—please verify before citing.