
Agent memory vs documentation: which slot a memory plugin fills, and which one the repo does
Tools in this post
Agent memory and documentation fill different slots. A memory plugin records what happened in past sessions, stores it outside the repository and decides, usually by similarity, what to put back into context. Documentation is written on purpose, checked into the repo, reviewed like code and loaded by a fixed rule. The argument reached the Hacker News front page on 3 October 2026 with Kevin Liao's post Agents don't need memory, they need documentation, at 369 points and 280 comments by 5 October. Neither side has a coding benchmark behind it.
Memory benchmarks measure chat recall, and the largest AGENTS.md study found no significant effect on task success. This post defines the two slots, sorts the tools into them, and follows the AGENTS.md vs CLAUDE.md vs skills post on what each file should hold.
What does an agent memory plugin do?
It captures notes from your sessions, stores them outside the repository and puts some of them back into context later. The pipelines differ more than Liao's post allows. His description, retrieve the five most similar snippets on every prompt and inject them, fits the Supermemory plugin. Claude Code auto memory and Codex Memories load Markdown at session start, Letta keeps its blocks in context permanently, and agentmemory ships with injection off.
| Tool | Stores in | Back into context | Token figure |
|---|---|---|---|
| Claude Code auto memory | ~/.claude/projects/<project>/memory/, Markdown, one machine |
MEMORY.md at session start, first 200 lines or 25KB |
A product limit |
| Codex Memories | ~/.codex/memories/ |
Later sessions, off by default | Not documented |
| Supermemory plugin | Supermemory's cloud | Each turn when Claude decides it helps, 5 profile items by default | None published |
| claude-mem | Local SQLite and Chroma | An index of about 50 observations at session start, MCP search on demand | About 800 tokens, vendor docs |
| Mem0 | Managed platform or open-source SDK | Top results injected per query | Under 7,000 per retrieval, vendor |
| Letta | Memory blocks | Always in context | Per-block character limits |
claude-mem has 96,523 GitHub stars, more than Mem0's 66,601. Every accuracy number in this space comes from a vendor running a chat-recall benchmark such as LoCoMo or LongMemEval, and the vendors contradict each other. Mem0's own 2025 paper shows full context beating Mem0 on answer quality, 72.90% against 66.88%. The memory layer won on cost: 1,764 tokens per query against 26,031.
An independent comparison, Zhou et al. in June 2026, tested 12 memory systems on conversational and long-context QA. Similarity retrieval fell from 37.1 to 7.4 F1 as the evidence got older, and append-only stores returned stale facts, which the authors call "hallucinations of the past". No coding tasks were in it.
Why is documentation a different slot from memory?
Because of who writes it, who reviews it and where it lives. Anthropic's memory docs draw the line: CLAUDE.md is "instructions you write", auto memory is "notes Claude writes itself", stored on one machine and not shared. OpenAI's Codex docs call memories "a helpful recall layer, not as the only source for rules that must always apply". Documentation ships with the code and loads by rule.
| Memory slot | Documentation slot | |
|---|---|---|
| Written by | The agent or a plugin hook | You, or the agent through review |
| Stored in | Home directory, SQLite, a vendor's cloud | The repository |
| Loaded by | Similarity, recency or the agent's choice | A rule: session start, file path, invocation |
| Reviewed in | Nowhere by default | Pull requests |
| Shared with | One machine, one user | Everyone who clones |
| Undo | Delete a note, if you find it | git revert |
Neither slot enforces anything. Claude Code's docs treat both CLAUDE.md and auto memory as "context, not enforced configuration", and a rule that must hold belongs in a PreToolUse hook.
Three projects have already gone from a store to plain files or search. Cursor removed Memories in version 2.1 in November 2025 and told users to export them into Rules. repository-harness ended its SQLite protocol on 10 August 2026. Boris Cherny, who created Claude Code, wrote on X in February that early versions used "RAG + a local vector db" before the team found "agentic search generally works better". That is his account, with no numbers.
What does a Git-versioned repository index give an agent that RAG does not?
A description that exists before the question. RAG answers "which stored chunks look like this prompt". A versioned index such as aoci-code states what each file is for, what callers rely on and what must not change, and you can diff, review and roll it back with Git. The cost is size: aoci-code reports an index of about 300K tokens for a 700,000-line system. Nobody has tested it independently.
aoci-code is a local Go MCP server. Your coding agent reads every managed file and writes one line per file into plain-text index files in the repo, with four fields:
| Field | Holds |
|---|---|
| F | The file's responsibility |
| R | What to read with it |
| A | What callers depend on |
| S | "What you cannot infer from the code but must not get wrong" |
The agent reads the whole index at task start in 7,000-token chunks, and the server flags entries that drifted from the code. The project's own estimate for building it is about an hour of agent time per 200,000 lines.
| aoci-code | DOX | repository-harness | |
|---|---|---|---|
| What ships | MCP server, CLI, index files | One Markdown instruction for a tree of AGENTS.md files | AGENTS.md, a docs map, templates, opt-in skills |
| Stars, 5 October | 1,151 | 1,478 | 1,234 |
| License | FSL-1.1-MIT, source-available | MIT | MIT |
| Releases | 18 release candidates, no stable | None, 6 commits | harness-v0.1.10, 13 August |
| Runtime | A local server | None | None after install |
The evidence on context files cuts against the overview part of an index. The ETH Zurich study (v3, 29 September 2026) found that "repository overviews, although popular and recommended by model providers, are not helpful". LLM-written context files changed success by -0.5% and -2%, and developer-written ones by +2.4%, none of it statistically significant, while cost rose up to 23%. Much of the coverage still quotes the v1 figures of +4% and -3%. Khatri (July 2026, 17 tasks, 288 runs) found no measurable correctness change from always-on or on-demand context. The field that might pay is S, because the code cannot supply it.
Which tools fill the memory slot and which fill the documentation slot?
Sort them by three properties, not by what they call themselves. aoci-code describes itself as "memory" on GitHub and is a documentation-slot tool. The test:
stored in the repo + changed through review + loaded by rule -> documentation slot
stored off the repo + written by the agent + loaded by score -> memory slot
anything else -> hybrid, check each property
- Memory slot: Claude Code auto memory, Codex Memories, Supermemory, claude-mem, Mem0 and OpenMemory, Zep, Letta, agentmemory.
- Documentation slot: AGENTS.md, CLAUDE.md, Cursor rules, skills, aoci-code, DOX, repository-harness, and Carson Gross's Markdown in /src proposal.
- Hybrids: Liao's own Operator Memory keeps
.operator-shared/in the repo and.operator/local. rag-rat anchors memories to source lines but keeps them in a SQLite file per machine.
Liao has a stake in the documentation side, because the post promotes his own tool. Operator Memory has no benchmark, which he confirmed on 4 October, and its public repo dates from 16 August although his post says he has used the system "for over a year". On HN, CapitalistCartr turned his argument back on him: "agents can't search for what they don't know. A markdown 'brain' has the same problem."
Where does this sit next to AGENTS.md, CLAUDE.md and skills?
AGENTS.md and CLAUDE.md are the documentation slot's always-loaded tier, skills its on-demand tier, and auto memory sits beside them as a notebook on one machine. In Claude Code, the project CLAUDE.md and MEMORY.md are re-read from disk after compaction, so both survive a context reset. A memory plugin survives only if it injects again. Instructions given only in chat do not survive at all.
The load path in Claude Code, from its memory, skills and context window docs, with the plugin and index rows added:
session start CLAUDE.md, or AGENTS.md when no CLAUDE.md exists in full, under 200 lines advised
MEMORY.md first 200 lines or 25KB
skill names and descriptions about 1% of the window
each prompt memory plugin recall top items by score
on file access nested CLAUDE.md, path-scoped rules when a matching file is read
on demand skill body when invoked
repo index over MCP (aoci-code) 7,000-token chunks
after compact CLAUDE.md, MEMORY.md re-read from disk
invoked skill bodies 5,000 tokens each, 25,000 total
chat-only instructions summarised away
Other agents cap the always-loaded tier differently: OpenAI Codex reads up to 32 KiB of AGENTS.md by default, and Cursor advises keeping each rule under 500 lines. OpenCode reads AGENTS.md first and falls back to CLAUDE.md.
What do you write down, and what do you let the agent rediscover?
Write down what the code cannot show: why a constraint exists, environment facts such as seeded test accounts, conventions that differ from tool defaults, and mistakes the agent made twice. Let it rediscover structure, the directory layouts, dependency lists and architecture overviews that the ETH Zurich study found unhelpful. Turn rules that must always hold into hooks, tests or lint rules.
An AGENTS.md diff in that direction (illustrative):
- ## Project structure
- src/api/ HTTP handlers
- src/service/ business logic
- Use 2-space indentation.
+ ## Constraints
+ - Payment retries stop at 3: the processor flags the merchant account at 5.
+ - QA accounts come from `make seed-qa`. Their passwords are in the vault, not .env.
+ - Migrations are append-only. Production runs them before the new code ships.
The sources agree on the cut. Claude Code's /doctor "cuts content Claude can derive from the codebase, such as directory layouts, dependency lists, and architecture overviews". A study of 100 repositories found 62% of agent files restating rules a linter already enforces. Gross puts what is left as "what it does, why it does it, and what it must not do".
Memory still has a job: catching the second mistake before you write it down. The top-ranked HN reply, from kaydub, said "The code IS the documentation", and replies under it warned that LLM-written docs "never trim, only amend". Another Claude Code user, b-karl, prunes auto memory regularly and moves anything durable into skills and repo docs. Let the agent take notes, and promote the ones that still hold a week later into a reviewed file.
On Stackness, as of 5 October 2026, 8 real profiles list Claude Code, 8 list Cursor, 3 list OpenCode and 2 list Codex. None lists AGENTS.md, DOX, Letta or a memory plugin (data sources). The numbers are small, and the AI coding tools developers list on Stackness are still the agents, not the layers around them.
Key numbers
- 369 points and 280 comments on the Liao post's HN thread, as of 5 October 2026 (HN).
- +2.4% success from developer-written context files, -0.5% and -2% from LLM-written ones, none significant, with cost up to 23% (ETH Zurich, v3, 29 September 2026).
- 72.90% for full context against 66.88% for Mem0 on LoCoMo, at 26,031 against 1,764 tokens per query (Mem0 paper, April 2025).
- 37.1 to 7.4 F1: similarity retrieval as evidence ages, across 12 memory systems (Zhou et al., June 2026).
- 200 lines or 25KB: what Claude Code loads of
MEMORY.mdper session (docs). - About 300K tokens: aoci-code's self-reported index for a 700,000-line system, 5 October 2026.
- 0 real Stackness profiles list AGENTS.md or a memory plugin, as of 5 October 2026 (data sources).
Quick answers
Agent memory vs documentation: what is the difference? Memory is what an agent or plugin recorded from past sessions, stored outside the repo and replayed by similarity or at session start. Documentation is written on purpose, committed, reviewed in pull requests and loaded by a fixed rule.
Do coding agents need a memory plugin? No independent study has measured one on coding tasks. Memory benchmarks test chat recall, and the vendors run them. A plugin can still catch repeated mistakes, which you then promote into AGENTS.md or CLAUDE.md.
Does AGENTS.md improve coding agent results? Not measurably in the largest study: developer-written files gave +2.4% and LLM-written ones -0.5% to -2%, none statistically significant, with cost up to 23%. Short files that hold what the code cannot show are the defensible version.
What is aoci-code? A source-available, Git-versioned index of a codebase, written by your coding agent and served by a local MCP server. It has shipped release candidates only, and its numbers are self-reported.
Does memory survive context compaction in Claude Code? CLAUDE.md and the first 200 lines of MEMORY.md are re-read from disk after compaction. Instructions given only in chat are summarised away.
Tools in this post
Use any of these tools?
Put them on a Stackness profile, say how you use each one and see who pairs them the same way. It takes a couple of minutes.


