~/satyajit

Agent Memory Repo: Devin's dreaming memory is a git repo of one-line bullets

mdjsonmcp

2026-10-06 · 20 min · agents · agent-memory · context-management · agentic-coding

Why read this

Essentialtop 10%

Reads the whole Agent Memory Repo spec and runs six concurrency cases in real git: appends conflict, and a stale session can quietly bring back a deleted fact.

  • Original, source-checked analysis
  • Runs on a laptop CPU
  • A lasting reference

Agents & harnessesMITPractitioner tool

How this was scored
Is it new?
1 of 3: An incremental tweak
Can I trust it?
3 of 3: Reproduces the headline result, or shows from primary files it is wrong
Can I run it?
3 of 3: Open, permissive, runs on reader hardware with instructions
Will I understand it?
2 of 3: Mechanism from first principles with figures
Can I act on it?
2 of 3: A concrete recipe, numbers or comparison
Will it last?
2 of 3: A reference for a year or more
Does it affect many?
1 of 3: A specialist community
Only here?
3 of 3: The only place this analysis exists

Score 76 of 100, ranked 44 of 445 rated articles. Each question is answered 0–3 by hand, and a 3 is rare. How articles are scored

On 5 October 2026 Cognition shipped two things in Devin. Memory: across sessions, Devin writes down what it learns about how you work. Dreaming: once a day a background session rereads those notes, merges duplicates, drops stale ones and writes down lessons that no single session stated. Walden Yan's post on X added the part that is useful outside Devin. The format is open, and it is called Agent Memory Repo: "support graph relationships, updates over time with historical records, backed by git and markdown, and you can use it with any agent."

That is three claims about a format, so I read the format. The spec repository, AgentMemoryRepo/agentmemoryrepo, holds five files: SPEC.md, README.md, one agent skill, a plugin manifest and an MIT licence. SPEC.md is 2,826 bytes (measured). The repo has 11 commits, all dated 4 and 5 October 2026 (measured, git log on the clone).

AgentMemoryRepo/agentmemoryrepo@1db04a5 · snapshot 2026-10-06
tracked files
5
license
MIT
branch
main
tests
none found
commit date
2026-10-05

local clone, 2026-10-06 at 1db04a5 — branch, commit, commitDate, fileCount, hasTests, license, licenseFile, shallow

shallow clone: counts describe the pinned tree, not the history

I read all of it, plus the cognition.com page, the Devin launch post and both X threads with their replies. Then I built a spec-shaped repo in a scratch directory and ran the concurrency cases through real git.

Short version. The format is small, and it is real. The "graph" is wikilinks between files. The "historical records" are git log. Nothing in the format itself says when a fact stopped being true. Dreaming, the part that makes it more than a folder of notes, is described but not released.

Why memory between sessions is hard

A coding agent starts every session knowing nothing about you. You tell it to use bun, not npm. You tell it that schema changes ship in their own PR. Next session it has forgotten both. The obvious fix, a rules file the agent always reads, works until the file is long. Every line costs context in every session, and nobody prunes it.

Any memory system has to answer three questions:

  1. What does it store? Raw transcripts, extracted facts, summaries, or rules.
  2. How does a session find the relevant part? Load all of it, retrieve by embedding, or navigate.
  3. How does a fact change? Overwrite it, delete it, or mark it no longer true and keep it.

Agent Memory Repo gives plain answers. It stores short Markdown facts. A session loads one small index and navigates from there with grep and links. A fact changes by editing the file, and git keeps the old version.

The format, concretely

A memory repo is a git repository. Its root is the memory root. Layout is free: the spec allows "Markdown notes, SQL queries, scripts, and other files". This is the spec's example:

memory-joe/              ← memory root
  MEMORY.md              ← entry point (required)
  team_structure.md
  using_datadog_mcp.md
  projects/
    payments.md
    website.md
  billing/
    count_paying_customers.sql

MEMORY.md is the one required file. "Agents load it at the start of every session." The spec says to keep it short. Entries every session needs go at the top. Links to everything else go under an ## Index heading:

# Memory: Joe
 
- Joe leads the product team [source: https://example.com/sessions/100]
 
## Index
- [[team_structure]]
- [[projects/payments]]
- [[projects/website]]

An entry is one bullet on one line, with optional metadata at the end. Metadata is [key: value; key: value]. Keys are open. Two are recommended: source, a link to the agent session where the fact was learned, and added, the date it was saved as YYYY-MM-DD.

- Joe coordinates the billing launch [source: https://example.com/sessions/101]
- Payments and website share a 2026-10-15 launch deadline [source: https://example.com/sessions/102; added: 2026-09-03]

Cross-links are [[path]]. Paths start at the memory root. You omit .md for Markdown and keep any other extension, so [[billing/count_paying_customers.sql]] points at a saved query. "Keep information in one place and link to it elsewhere." When a file moves, the agent updates the links.

Several repos can share one session. Alice's memory is cloned at session start. When Bob joins, his memory is cloned beside it, only if he chooses to share it. Each repo keeps its own MEMORY.md, permissions and history. A [[path]] resolves from the root of the repo that contains it. The agent writes Alice's preferences to memory-alice/ and Bob's to memory-bob/, and it asks when the destination is unclear.

That is the whole spec. There is no YAML frontmatter anywhere in it: no type, no description, no schema version. That matters for the comparison later, because both Claude Code and Letta put frontmatter on every memory file.

The widget reads one line the way the spec describes. The presets come from SPEC.md, the README and the cognition.com page, plus one edge case I made up. Three of them hit places where the four-sentence grammar leaves the reader to decide.

one entry, read by the specfrom: SPEC.md
factPayments and website share a 2026-10-15 launch deadlinemetadata
source = https://example.com/sessions/102added = 2026-09-03
links (edges)
none

The page's own swarm example writes an entry with three separate [source: …] brackets. SPEC.md describes one bracketed list. The same example nests replies as indented sub-bullets ("Cache answers: yes…"), and the spec never says whether an indented bullet is an entry. The ; separator has no escape, so a URL that contains ; is cut in two. None of this breaks an LLM reading the file; a model reads around it. It does mean two tools that both claim to parse the format can disagree on the same line. A reference parser or a conformance test would settle it, and the repo has neither (measured: the clone holds Markdown, a JSON manifest and a licence, and no parser or test code).

The graph is the set of [[path]] links. A node is a file, not an entry, and an edge has no type. [[team_structure]] says "see that file". It cannot say "reports to", "replaced by" or "contradicts". The launch video shows exactly this: after dreaming, notes in different folders are joined by plain curves.

A frame from Cognition's launch video: a memory/ folder tree with bun.md, redis.md, rate-limits.md, migrations.md and staging.md joined by curved link lines on the left, and a new highlighted file conventions.md reading 'Prefer existing tools: bun, lib/redis, middleware/.' Caption text: 'And new knowledge emerges.'
Dreaming as Cognition's launch video draws it: notes linked across folders, then a new note, conventions.md, written from what the others share. An animation, not a recording of a run (Cognition, launch video on X).

That is a weaker graph than "knowledge graph" usually means, and it is a deliberate trade. A typed graph needs a schema and an extractor that fills it. A file graph needs an agent that can write [[...]], which every agent can. Anyone who keeps an Obsidian vault will recognise the shape. The reply asking "How is this different from Obsidian" is fair. The answer is the root-relative paths, the one-line entries with source links, and git.

History lives in git, not in the format

The spec's whole rule for changing a fact is one line: "Update or remove entries when the information changes." The skill makes it stricter: "Edit files in place. Update or remove entries that are out of date instead of adding a contradicting entry."

So the working tree is always the current belief, and only the current belief. There is no superseded_by, no valid_until, no removed key. The only time-like key is added. "Updates over time with historical records" means the record lives in git. Asking when the dev-server port note came and went is a pickaxe search:

$ git log --format='%h %s' -S'port 3001'
535bfb9 dream: drop dev-port
f9bd7f0 remember port
371bad7 init

That output is from my scratch repo (measured). It is a good design for one reason: the history costs nothing at session start. The agent loads a small current index and digs into history only when it asks. It has a cost too. An agent that reads only the working tree cannot tell "never knew this" from "knew this and deliberately forgot it." One reply under Walden Yan's post asked exactly that: "Does the format have tombstones for deleted memories? Otherwise an agent importing an older clone can bring back something the user explicitly asked to forget." It does not. The next section shows when that matters.

Concurrency: the memory loop under real git

The spec's memory loop is four steps: clone the latest memory, grep or follow links, update entries "with no human in the loop", commit after every edit. The cognition.com page phrases the steps as clone, search, update, push. Devin's own drive adds a revision check, shown in the launch post's interactive figure: a session whose checkout is behind is refused, merges, and saves again.

Devin's 'Two sessions saving to one memory drive', step 4 of 6. Session A has saved backend/redis.md and is at rev 2. Session B, still at rev 1 with an edited tooling/bun.md, tries to save and is marked rejected and stale. The memory drive is at rev 2, updated by A.
Devin's revision check: Session B's save is refused because its copy is based on revision 1. In the next step B merges revision 2 and saves again (Cognition, Devin launch post, interactive figure step 4 of 6).

The cognition.com page describes the swarm case like this: "Git merges most of their commits. When two agents edit the same line, git rejects the second push, and that agent reads both versions before writing one." That is close, but not what git does. Git rejects every push from a checkout that is behind, whatever lines it touched. The merge that follows is where lines matter, and the conflict condition is wider than "the same line". I ran six cases in a bare repo with two clones, letting session A push first and having B commit from the older checkout and pull (measured, git 2.43.0):

CaseResult of B's git pull --no-rebase
1. A and B each append one bullet to the end of findings.mdCONFLICT
2. Same, with findings.md merge=union in .gitattributesclean, both lines kept
3. A and B edit different filesclean
4. A (dreaming) deletes an entry; B never touched itclean, entry stays deleted
5. A deletes an entry; B edits that lineCONFLICT
6. A deletes an entry; B writes the same fact into a new fileclean, fact is back

In cases 1 to 3 I also had B push before pulling, and git refused the push as non-fast-forward every time. That includes case 3, where nothing overlaps.

two sessions, one memory repo, plain gitmeasured, git 2.43.0
file: findings.md · common ancestor:
- Database: queries are fast [source: s/301]
A · runtime agent (pushed first)
- Database: queries are fast [source: s/301]
- Runtime: GC freezes 400 ms [source: s/302]
B · load-balancer agent
- Database: queries are fast [source: s/301]
- LB: ruled out [source: s/303]
B is behind, so git push is refused (non-fast-forward) · git pull --no-rebase →CONFLICT
- Database: queries are fast [source: s/301]
<<<<<<< HEAD
- LB: ruled out [source: s/303]
=======
- Runtime: GC freezes 400 ms [source: s/302]
>>>>>>> (the commit B pulled)

Both sides inserted a line at the same place, the end of the file. Git cannot order two insertions at one point, so a shared findings file conflicts on every pair of concurrent appends, not only when two agents edit the same line.

Each run starts from the same two files, lets session A push first, then has B commit from the older checkout and pull. The outcomes are what git printed; the widget only shows them side by side.

Case 1 matters most for the swarm example, where four agents "write to the same two files the whole time". Two appends at the end of a file insert at the same point, and git refuses to choose an order. Every pair of concurrent appends conflicts. The agent can resolve it, since it is two bullets and both are kept, but that is an LLM round-trip per collision. One line of .gitattributes removes it for log-like files (case 2). One file per agent, as the swarm example's agents/ folder already does, removes it too (case 3). The spec mentions neither.

Cases 4 to 6 answer the tombstone question. A stale clone does not bring a deleted fact back by merging (case 4). Three-way merge sees that it never changed the line. An edit to the deleted line is a real conflict (case 5), and the spec gives the resolving agent nothing to decide with. The bad case is 6. An older session re-learns the fact from its own stale context and writes it somewhere new. Git sees a new line in a new file, and the fact quietly returns. Only a tombstone, an entry that says "removed on purpose, on this date, because of this session", could catch it. Dreaming could catch it too, if it checks new entries against past deletions in git log. Neither is specified.

Dreaming: what the posts say it does

The spec gives dreaming two sentences: a dedicated agent that runs periodically, to "add new memory" (spot patterns across sessions) and "clean up memory" (merge duplicates, remove outdated entries, and check sources to resolve contradictions). The Devin launch post is more specific. Dreaming is "a daily background session" in which Devin "reviews past conversations alongside its existing memory". It:

The post's second interactive figure steps through one pass. Seven flat notes from a week of sessions go in. landing-page.md and styling.md describe the same web stack, so they merge into stack.md, and the links to both source sessions are kept. dev-port.md ("Dev server was on port 3001 this afternoon") is unused and transient, so it is dropped. The rest are regrouped into tooling/, frontend/ and backend/, and the index is rewritten. Then two notes about ordering produce a third that neither states: ship in small PRs, schema first, then API, then UI.

Devin's 'One dreaming pass over the memory drive', step 6 of 6: a memory/ tree with tooling/bun.md, frontend/stack.md, frontend/components.md, backend/migrations.md and backend/api.md, the last two tagged 'source', and a new workflow/conventions.md tagged 'new' reading 'Ship in small PRs: schema first, then API, then UI.'
The end of one dreaming pass: notes regrouped by topic, and a new note inferred from two others, linked to the sessions it came from (Cognition, Devin launch post, interactive figure step 6 of 6).

In the product, each dreaming run shows up as a card under Customize → Memory, with a one-line summary and a diff size. The demo clip shows runs such as "Merged duplicate tooling notes and linked stack.md to conventions.md", with three files changed.

Devin's Customize → Memory page: a 'Personal' memory described as 'Personal to you and not shared with your organization', three 'Recent dreaming sessions' cards (Refinement 8h ago: merged duplicate tooling notes, 3 files +9 −11; 1d ago: categorized backend learnings, 3 files +27 −7; 2d ago: captured the schema-PR-first rule, 2 files +8 −1), and a file tree with MEMORY.md open under 'How I work with you'.
Dreaming runs as the Devin UI lists them, each a reviewable diff on the memory drive. A frame from a product demo, in an example organization called Acme (Cognition, Devin launch post, demo video).

Three things are worth stating plainly.

The cadence is described three ways. "At night" (Cognition's post), "a daily asynchronous process" (the launch blog), "runs periodically" (the spec). They are compatible. Only the spec's wording binds anyone else.

Dreaming is not in the open repo. The skill says so: "This skill does work only when it is invoked. It doesn't add hooks, scheduled jobs, or startup scripts." The cognition.com trial ends with "Automatic startup and scheduled Dreaming are not included." The open part is the format and a skill for reading and writing it. The consolidation prompt, the "unused by any session" signal and the scheduler stay inside Devin.

There are no numbers. Neither post reports how much dreaming shrinks memory, whether it improves task success, or how often it deletes something a user wanted kept. Cognition's thread says a swarm of Devins "using the memory as a way to coordinate at large scale without us prompting them" appeared in early experiments, with no data attached. The one number in the replies is a user's anecdote. Claude Code kept a project's memory for 10 days, then a fresh Devin session dreamed on it: "21 commits, 19 → 14 notes, duplicates merged" (reported, a reply on X, not checked). It is a plausible demo of the format crossing agents. It is not an evaluation.

How it compares

None of this is new as an idea. What is new is a written, MIT-licensed spec with a plain-text format.

Claude Code's auto memory is the nearest sibling, and the shapes are close. Each project gets ~/.claude/projects/<project>/memory/ with a MEMORY.md index, "one line per memory, loaded into every session", plus one topic file per memory, read on demand. The first 200 lines or 25 KB of the index load at session start (reported, Claude Code docs). The differences: each memory file carries a type field in frontmatter (user, feedback, project, reference), the directory is "machine-local" with no git and no sharing, and the docs describe no scheduled consolidation. Claude Code nudges the model to merge or drop stale entries when the index nears its limit.

Table from the Claude Code memory documentation comparing CLAUDE.md files and auto memory: who writes it (you vs Claude), what it contains (instructions and rules vs learnings and patterns), scope (project, user, or org vs per repository, shared across worktrees), loaded into (every session vs every session, first 200 lines or 25KB), use for.
Claude Code's two memory systems side by side. Auto memory uses the same MEMORY.md-plus-topic-files shape, kept on one machine (Claude Code documentation, Memory page).

Letta got here first on almost every axis. MemGPT (arXiv 2310.08560) framed memory as an operating system's paging problem. A fixed context window holds system instructions, a read-write working context and a FIFO message queue, and the model moves data to and from archival and recall storage through function calls.

MemGPT system diagram: the LLM finite context window split into System Instructions (read-only), Working Context (read-write via functions) and a FIFO Queue (read-write via queue manager), feeding an Output Buffer; below, Archival Storage and Recall Storage connect through a Function Executor and a Queue Manager.
MemGPT's memory hierarchy: the model pages facts between its context and external stores with function calls (MemGPT paper, Figure 3).

In February 2026 Letta rebuilt Letta Code's memory as Context Repositories. It is a git-backed memory filesystem cloned locally, "every change to memory is automatically versioned with informative commit messages", and subagents work in their own git worktrees and merge back. It has a background "sleep-time" reflection process and a defragmentation skill that merges duplicates and restructures memory into "a clean hierarchy of 15–25 focused files" (reported, Letta blog). Two replies under Cognition's post made the same point, one of them bluntly. Letta's files carry a frontmatter description, and a system/ folder marks what is always loaded.

A directory tree from Letta's context repositories post: system/ containing human/prefs/coding_style.md, communication.md, workflow.md, human/identity.md, project/gotchas.md, project/overview.md and persona.md; history/ containing project_timeline.md, claude.md and codex.md; and api_documentation.md at the root.
A Letta context repository: files under system/ are pinned into the prompt, the rest are read on demand (Letta, Context Repositories blog post).

The honest comparison: Agent Memory Repo is close to Letta's context repositories with less structure. It has no frontmatter and no pinned folder, and it adds [source:] provenance per line and the multi-repo composition story. Letta's idea that sleep-time compute can do useful work between requests has its own paper (arXiv 2504.13171).

Zep's Graphiti sits at the other end. It is a typed, temporal knowledge graph (arXiv 2501.13956). Every fact is an edge with four timestamps: when the system created and expired it, and the interval during which it "held true". When a new edge contradicts an old one, Graphiti does not delete anything. It "invalidates the affected edges by setting their tinvalidt_{\text{invalid}} to the tvalidt_{\text{valid}} of the invalidating edge." That is supersession inside the data, where any query can see it, which is exactly what Agent Memory Repo leaves to git log.

Graphiti's README animation, final frame: a Kendra node linked to Puma Shoes by 'likes' (valid_at 2024-09-03T13:15:22Z) and to Adidas Shoes by 'broke' (valid_at 2024-09-03T13:15:22Z) and by 'loves', which now carries invalid_at 2024-09-03T13:15:22Z.
Supersession in a temporal graph: the old 'loves' edge is kept and stamped invalid_at rather than deleted (Graphiti README animation, final frame).

Mem0 (arXiv 2504.19413) is the extraction-pipeline version. An LLM extracts candidate facts from each exchange, retrieves the most similar stored memories, and picks one of four operations: ADD, UPDATE, DELETE or NOOP. In base Mem0, DELETE removes the contradicted memory. The graph variant marks conflicting relationships "as invalid rather than physically removing them to enable temporal reasoning."

Mem0 architecture: messages enter an Extraction Phase (an LLM with a summary and the last m messages) producing new extracted memories; an Update Phase fetches the top s similar memories from a database and a tool call chooses ADD, UPDATE, DELETE or NOOP, writing new memories back to the database.
Mem0's write path: extract facts, compare with similar memories, then add, update, delete or do nothing (Mem0 paper, Figure 2).
StoreUnitLinksWhen a fact changesConsolidationOpen
Agent Memory Repogit + Markdownone-line bullet[[path]], file to file, untypededit or delete; old text in git logdreaming, described, not releasedspec + skill, MIT
Claude Code auto memorylocal Markdownone file per memory, type frontmatterindex linesedit or deletenudged at the index limitdocs
Letta context repositoriesgit + Markdownfile with description frontmatterfolder treeedit; git historysleep-time reflection, defragLetta Code
Zep / Graphitigraph databasetyped edge with 4 timestampstyped edgesold edge stamped invalidentity resolution, community summariesGraphiti open source
Mem0vector store (+ graph)extracted factgraph variant onlyUPDATE / DELETE; graph marks invalidon every writeopen source

Every row is reported from the cited source. I did not run Graphiti, Mem0 or Letta for this piece.

Checking the claims

What I would do with it

I would use the format. It is small enough to hold in your head. A source link on every fact is the right default, and memory you can git revert beats memory you cannot see.

I would add three things the spec leaves out. Put merge=union on any append-only file, or give each writer its own file. Add a removed entry instead of deleting what a user asked to forget, so dreaming and other sessions can see the decision. And add scope as a metadata key, since "this repo", "this task" and "every project" are different facts. Keys are open, so all three fit in today's spec. Whether two tools would agree on them is the open question an open standard is supposed to answer. With 11 commits and no parser in the repo, it has not answered it yet.

Related on this site: Hindsight organises memory into four typed networks and benchmarks it. TencentDB Agent Memory sorts its tiers by how often they change. MemHarness argues retrieved memory should be rewritten before use, not replayed. Agent harnesses covers why durable agent state belongs in files. And Dream-RSI is a different "dreaming": a loop that rewrites a search scheduler, not a memory.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "Agent Memory Repo: Devin's dreaming memory is a git repo of one-line bullets", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026agentmemoryrepo,
  author = {Satyajit Ghana},
  title  = {Agent Memory Repo: Devin's dreaming memory is a git repo of one-line bullets},
  url    = {https://ai.thesatyajit.com/articles/agent-memory-repo},
  year   = {2026}
}
share