2026-10-06 · 20 min · agents · agent-memory · context-management · agentic-coding
Why read this
Essentialtop 10%Reads the whole Agent Memory Repo spec and runs six concurrency cases in real git: appends conflict, and a stale session can quietly bring back a deleted fact.
- Original, source-checked analysis
- Runs on a laptop CPU
- A lasting reference
Agents & harnessesMITPractitioner tool
How this was scored
- Is it new?
- 1 of 3: An incremental tweak
- Can I trust it?
- 3 of 3: Reproduces the headline result, or shows from primary files it is wrong
- Can I run it?
- 3 of 3: Open, permissive, runs on reader hardware with instructions
- Will I understand it?
- 2 of 3: Mechanism from first principles with figures
- Can I act on it?
- 2 of 3: A concrete recipe, numbers or comparison
- Will it last?
- 2 of 3: A reference for a year or more
- Does it affect many?
- 1 of 3: A specialist community
- Only here?
- 3 of 3: The only place this analysis exists
Score 76 of 100, ranked 44 of 445 rated articles. Each question is answered 0–3 by hand, and a 3 is rare. How articles are scored
On 5 October 2026 Cognition shipped two things in Devin. Memory: across sessions, Devin writes down what it learns about how you work. Dreaming: once a day a background session rereads those notes, merges duplicates, drops stale ones and writes down lessons that no single session stated. Walden Yan's post on X added the part that is useful outside Devin. The format is open, and it is called Agent Memory Repo: "support graph relationships, updates over time with historical records, backed by git and markdown, and you can use it with any agent."
That is three claims about a format, so I read the format. The spec repository, AgentMemoryRepo/agentmemoryrepo, holds five files: SPEC.md, README.md, one agent skill, a plugin manifest and an MIT licence. SPEC.md is 2,826 bytes (measured). The repo has 11 commits, all dated 4 and 5 October 2026 (measured, git log on the clone).
- license
- MIT
- branch
- main
- tests
- none found
- commit date
- 2026-10-05
local clone, 2026-10-06 at 1db04a5 — branch, commit, commitDate, fileCount, hasTests, license, licenseFile, shallow
shallow clone: counts describe the pinned tree, not the history
I read all of it, plus the cognition.com page, the Devin launch post and both X threads with their replies. Then I built a spec-shaped repo in a scratch directory and ran the concurrency cases through real git.
Short version. The format is small, and it is real. The "graph" is wikilinks between files. The "historical records" are git log. Nothing in the format itself says when a fact stopped being true. Dreaming, the part that makes it more than a folder of notes, is described but not released.
Why memory between sessions is hard
A coding agent starts every session knowing nothing about you. You tell it to use bun, not npm. You tell it that schema changes ship in their own PR. Next session it has forgotten both. The obvious fix, a rules file the agent always reads, works until the file is long. Every line costs context in every session, and nobody prunes it.
Any memory system has to answer three questions:
- What does it store? Raw transcripts, extracted facts, summaries, or rules.
- How does a session find the relevant part? Load all of it, retrieve by embedding, or navigate.
- How does a fact change? Overwrite it, delete it, or mark it no longer true and keep it.
Agent Memory Repo gives plain answers. It stores short Markdown facts. A session loads one small index and navigates from there with grep and links. A fact changes by editing the file, and git keeps the old version.
The format, concretely
A memory repo is a git repository. Its root is the memory root. Layout is free: the spec allows "Markdown notes, SQL queries, scripts, and other files". This is the spec's example:
memory-joe/ ← memory root
MEMORY.md ← entry point (required)
team_structure.md
using_datadog_mcp.md
projects/
payments.md
website.md
billing/
count_paying_customers.sqlMEMORY.md is the one required file. "Agents load it at the start of every session." The spec says to keep it short. Entries every session needs go at the top. Links to everything else go under an ## Index heading:
# Memory: Joe
- Joe leads the product team [source: https://example.com/sessions/100]
## Index
- [[team_structure]]
- [[projects/payments]]
- [[projects/website]]An entry is one bullet on one line, with optional metadata at the end. Metadata is [key: value; key: value]. Keys are open. Two are recommended: source, a link to the agent session where the fact was learned, and added, the date it was saved as YYYY-MM-DD.
- Joe coordinates the billing launch [source: https://example.com/sessions/101]
- Payments and website share a 2026-10-15 launch deadline [source: https://example.com/sessions/102; added: 2026-09-03]Cross-links are [[path]]. Paths start at the memory root. You omit .md for Markdown and keep any other extension, so [[billing/count_paying_customers.sql]] points at a saved query. "Keep information in one place and link to it elsewhere." When a file moves, the agent updates the links.
Several repos can share one session. Alice's memory is cloned at session start. When Bob joins, his memory is cloned beside it, only if he chooses to share it. Each repo keeps its own MEMORY.md, permissions and history. A [[path]] resolves from the root of the repo that contains it. The agent writes Alice's preferences to memory-alice/ and Bob's to memory-bob/, and it asks when the destination is unclear.
That is the whole spec. There is no YAML frontmatter anywhere in it: no type, no description, no schema version. That matters for the comparison later, because both Claude Code and Letta put frontmatter on every memory file.
The widget reads one line the way the spec describes. The presets come from SPEC.md, the README and the cognition.com page, plus one edge case I made up. Three of them hit places where the four-sentence grammar leaves the reader to decide.
The page's own swarm example writes an entry with three separate [source: …] brackets. SPEC.md describes one bracketed list. The same example nests replies as indented sub-bullets ("Cache answers: yes…"), and the spec never says whether an indented bullet is an entry. The ; separator has no escape, so a URL that contains ; is cut in two. None of this breaks an LLM reading the file; a model reads around it. It does mean two tools that both claim to parse the format can disagree on the same line. A reference parser or a conformance test would settle it, and the repo has neither (measured: the clone holds Markdown, a JSON manifest and a licence, and no parser or test code).
"Graph relationships" are links between files
The graph is the set of [[path]] links. A node is a file, not an entry, and an edge has no type. [[team_structure]] says "see that file". It cannot say "reports to", "replaced by" or "contradicts". The launch video shows exactly this: after dreaming, notes in different folders are joined by plain curves.

That is a weaker graph than "knowledge graph" usually means, and it is a deliberate trade. A typed graph needs a schema and an extractor that fills it. A file graph needs an agent that can write [[...]], which every agent can. Anyone who keeps an Obsidian vault will recognise the shape. The reply asking "How is this different from Obsidian" is fair. The answer is the root-relative paths, the one-line entries with source links, and git.
History lives in git, not in the format
The spec's whole rule for changing a fact is one line: "Update or remove entries when the information changes." The skill makes it stricter: "Edit files in place. Update or remove entries that are out of date instead of adding a contradicting entry."
So the working tree is always the current belief, and only the current belief. There is no superseded_by, no valid_until, no removed key. The only time-like key is added. "Updates over time with historical records" means the record lives in git. Asking when the dev-server port note came and went is a pickaxe search:
$ git log --format='%h %s' -S'port 3001'
535bfb9 dream: drop dev-port
f9bd7f0 remember port
371bad7 initThat output is from my scratch repo (measured). It is a good design for one reason: the history costs nothing at session start. The agent loads a small current index and digs into history only when it asks. It has a cost too. An agent that reads only the working tree cannot tell "never knew this" from "knew this and deliberately forgot it." One reply under Walden Yan's post asked exactly that: "Does the format have tombstones for deleted memories? Otherwise an agent importing an older clone can bring back something the user explicitly asked to forget." It does not. The next section shows when that matters.
Concurrency: the memory loop under real git
The spec's memory loop is four steps: clone the latest memory, grep or follow links, update entries "with no human in the loop", commit after every edit. The cognition.com page phrases the steps as clone, search, update, push. Devin's own drive adds a revision check, shown in the launch post's interactive figure: a session whose checkout is behind is refused, merges, and saves again.

The cognition.com page describes the swarm case like this: "Git merges most of their commits. When two agents edit the same line, git rejects the second push, and that agent reads both versions before writing one." That is close, but not what git does. Git rejects every push from a checkout that is behind, whatever lines it touched. The merge that follows is where lines matter, and the conflict condition is wider than "the same line". I ran six cases in a bare repo with two clones, letting session A push first and having B commit from the older checkout and pull (measured, git 2.43.0):
| Case | Result of B's git pull --no-rebase |
|---|---|
1. A and B each append one bullet to the end of findings.md | CONFLICT |
2. Same, with findings.md merge=union in .gitattributes | clean, both lines kept |
| 3. A and B edit different files | clean |
| 4. A (dreaming) deletes an entry; B never touched it | clean, entry stays deleted |
| 5. A deletes an entry; B edits that line | CONFLICT |
| 6. A deletes an entry; B writes the same fact into a new file | clean, fact is back |
In cases 1 to 3 I also had B push before pulling, and git refused the push as non-fast-forward every time. That includes case 3, where nothing overlaps.
- Database: queries are fast [source: s/301]
- Database: queries are fast [source: s/301]- Runtime: GC freezes 400 ms [source: s/302]
- Database: queries are fast [source: s/301]- LB: ruled out [source: s/303]
- Database: queries are fast [source: s/301]<<<<<<< HEAD- LB: ruled out [source: s/303]=======- Runtime: GC freezes 400 ms [source: s/302]>>>>>>> (the commit B pulled)
Both sides inserted a line at the same place, the end of the file. Git cannot order two insertions at one point, so a shared findings file conflicts on every pair of concurrent appends, not only when two agents edit the same line.
Case 1 matters most for the swarm example, where four agents "write to the same two files the whole time". Two appends at the end of a file insert at the same point, and git refuses to choose an order. Every pair of concurrent appends conflicts. The agent can resolve it, since it is two bullets and both are kept, but that is an LLM round-trip per collision. One line of .gitattributes removes it for log-like files (case 2). One file per agent, as the swarm example's agents/ folder already does, removes it too (case 3). The spec mentions neither.
Cases 4 to 6 answer the tombstone question. A stale clone does not bring a deleted fact back by merging (case 4). Three-way merge sees that it never changed the line. An edit to the deleted line is a real conflict (case 5), and the spec gives the resolving agent nothing to decide with. The bad case is 6. An older session re-learns the fact from its own stale context and writes it somewhere new. Git sees a new line in a new file, and the fact quietly returns. Only a tombstone, an entry that says "removed on purpose, on this date, because of this session", could catch it. Dreaming could catch it too, if it checks new entries against past deletions in git log. Neither is specified.
Dreaming: what the posts say it does
The spec gives dreaming two sentences: a dedicated agent that runs periodically, to "add new memory" (spot patterns across sessions) and "clean up memory" (merge duplicates, remove outdated entries, and check sources to resolve contradictions). The Devin launch post is more specific. Dreaming is "a daily background session" in which Devin "reviews past conversations alongside its existing memory". It:
- consolidates overlapping notes,
- removes transient details,
- looks for lessons that weren't captured during the original work,
- removes stale memory records that no session used.
The post's second interactive figure steps through one pass. Seven flat notes from a week of sessions go in. landing-page.md and styling.md describe the same web stack, so they merge into stack.md, and the links to both source sessions are kept. dev-port.md ("Dev server was on port 3001 this afternoon") is unused and transient, so it is dropped. The rest are regrouped into tooling/, frontend/ and backend/, and the index is rewritten. Then two notes about ordering produce a third that neither states: ship in small PRs, schema first, then API, then UI.

In the product, each dreaming run shows up as a card under Customize → Memory, with a one-line summary and a diff size. The demo clip shows runs such as "Merged duplicate tooling notes and linked stack.md to conventions.md", with three files changed.

Three things are worth stating plainly.
The cadence is described three ways. "At night" (Cognition's post), "a daily asynchronous process" (the launch blog), "runs periodically" (the spec). They are compatible. Only the spec's wording binds anyone else.
Dreaming is not in the open repo. The skill says so: "This skill does work only when it is invoked. It doesn't add hooks, scheduled jobs, or startup scripts." The cognition.com trial ends with "Automatic startup and scheduled Dreaming are not included." The open part is the format and a skill for reading and writing it. The consolidation prompt, the "unused by any session" signal and the scheduler stay inside Devin.
There are no numbers. Neither post reports how much dreaming shrinks memory, whether it improves task success, or how often it deletes something a user wanted kept. Cognition's thread says a swarm of Devins "using the memory as a way to coordinate at large scale without us prompting them" appeared in early experiments, with no data attached. The one number in the replies is a user's anecdote. Claude Code kept a project's memory for 10 days, then a fresh Devin session dreamed on it: "21 commits, 19 → 14 notes, duplicates merged" (reported, a reply on X, not checked). It is a plausible demo of the format crossing agents. It is not an evaluation.
How it compares
None of this is new as an idea. What is new is a written, MIT-licensed spec with a plain-text format.
Claude Code's auto memory is the nearest sibling, and the shapes are close. Each project gets ~/.claude/projects/<project>/memory/ with a MEMORY.md index, "one line per memory, loaded into every session", plus one topic file per memory, read on demand. The first 200 lines or 25 KB of the index load at session start (reported, Claude Code docs). The differences: each memory file carries a type field in frontmatter (user, feedback, project, reference), the directory is "machine-local" with no git and no sharing, and the docs describe no scheduled consolidation. Claude Code nudges the model to merge or drop stale entries when the index nears its limit.

Letta got here first on almost every axis. MemGPT (arXiv 2310.08560) framed memory as an operating system's paging problem. A fixed context window holds system instructions, a read-write working context and a FIFO message queue, and the model moves data to and from archival and recall storage through function calls.

In February 2026 Letta rebuilt Letta Code's memory as Context Repositories. It is a git-backed memory filesystem cloned locally, "every change to memory is automatically versioned with informative commit messages", and subagents work in their own git worktrees and merge back. It has a background "sleep-time" reflection process and a defragmentation skill that merges duplicates and restructures memory into "a clean hierarchy of 15–25 focused files" (reported, Letta blog). Two replies under Cognition's post made the same point, one of them bluntly. Letta's files carry a frontmatter description, and a system/ folder marks what is always loaded.

The honest comparison: Agent Memory Repo is close to Letta's context repositories with less structure. It has no frontmatter and no pinned folder, and it adds [source:] provenance per line and the multi-repo composition story. Letta's idea that sleep-time compute can do useful work between requests has its own paper (arXiv 2504.13171).
Zep's Graphiti sits at the other end. It is a typed, temporal knowledge graph (arXiv 2501.13956). Every fact is an edge with four timestamps: when the system created and expired it, and the interval during which it "held true". When a new edge contradicts an old one, Graphiti does not delete anything. It "invalidates the affected edges by setting their to the of the invalidating edge." That is supersession inside the data, where any query can see it, which is exactly what Agent Memory Repo leaves to git log.

Mem0 (arXiv 2504.19413) is the extraction-pipeline version. An LLM extracts candidate facts from each exchange, retrieves the most similar stored memories, and picks one of four operations: ADD, UPDATE, DELETE or NOOP. In base Mem0, DELETE removes the contradicted memory. The graph variant marks conflicting relationships "as invalid rather than physically removing them to enable temporal reasoning."

| Store | Unit | Links | When a fact changes | Consolidation | Open | |
|---|---|---|---|---|---|---|
| Agent Memory Repo | git + Markdown | one-line bullet | [[path]], file to file, untyped | edit or delete; old text in git log | dreaming, described, not released | spec + skill, MIT |
| Claude Code auto memory | local Markdown | one file per memory, type frontmatter | index lines | edit or delete | nudged at the index limit | docs |
| Letta context repositories | git + Markdown | file with description frontmatter | folder tree | edit; git history | sleep-time reflection, defrag | Letta Code |
| Zep / Graphiti | graph database | typed edge with 4 timestamps | typed edges | old edge stamped invalid | entity resolution, community summaries | Graphiti open source |
| Mem0 | vector store (+ graph) | extracted fact | graph variant only | UPDATE / DELETE; graph marks invalid | on every write | open source |
Every row is reported from the cited source. I did not run Graphiti, Mem0 or Letta for this piece.
Checking the claims
- "Support graph relationships." Holds, if a graph means untyped links between files. It is not a knowledge graph in Graphiti's sense, and per-entry links do not exist.
- "Updates over time with historical records." Holds through git, not through the format. The working tree has no record of what was superseded.
git log -Sdoes. - "Backed by git and markdown." Holds, and it is the strongest part. A memory change is a diff, which a person can review, revert or blame. The Devin UI shows each dreaming run as exactly that.
- "You can use it with any agent." Holds for reading and writing. The format is plain text, and the skill installs into Claude Code, Cursor and Devin. Dreaming is not portable, because it is not released.
- "Git flags any edits that conflict." Half true. Git also conflicts on edits that do not overlap, such as two appends at the end of one file (measured, case 1). It flags nothing when a stale session re-learns a deleted fact (measured, case 6).
What I would do with it
I would use the format. It is small enough to hold in your head. A source link on every fact is the right default, and memory you can git revert beats memory you cannot see.
I would add three things the spec leaves out. Put merge=union on any append-only file, or give each writer its own file. Add a removed entry instead of deleting what a user asked to forget, so dreaming and other sessions can see the decision. And add scope as a metadata key, since "this repo", "this task" and "every project" are different facts. Keys are open, so all three fit in today's spec. Whether two tools would agree on them is the open question an open standard is supposed to answer. With 11 commits and no parser in the repo, it has not answered it yet.
Related on this site: Hindsight organises memory into four typed networks and benchmarks it. TencentDB Agent Memory sorts its tiers by how often they change. MemHarness argues retrieved memory should be rewritten before use, not replayed. Agent harnesses covers why durable agent state belongs in files. And Dream-RSI is a different "dreaming": a loop that rewrites a search scheduler, not a memory.