2026-10-02 · 16 min · explainer · agents · agent-memory · retrieval · llm · benchmarks · open-source
A memory system trended to the top of GitHub, which does not happen often for a thing whose whole job is to remember what you said three sessions ago. Hindsight — from Vectorize — crossed 40,000 stars in days and is at about 44,420 as I write this (measured from the GitHub API, 2026-10-02). It is MIT-licensed Python, and it comes with a paper: Hindsight is 20/20: Building Agent Memory that Retains, Recalls, and Reflects (arXiv 2512.12818, submitted 2025-12-14).
- license
- MIT
- branch
- main
- tests
- 1206 files
- source
- 32.8 MB
- commit date
- 2026-10-01
by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded
local clone, 2026-10-02 at 0be6c02 — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow, testFileCount
shallow clone: counts describe the pinned tree, not the history
The star count is the news; the paper is the substance, and it contains one number that is worth the whole read. On LongMemEval, the same open-source 20B model scores 39.0% with the full conversation pasted into its context, and 83.6% when Hindsight sits in front of it (reported, paper Table 3). Same backbone, same judge, same questions — a 44.6-point swing that comes entirely from how the memory is organized, not from a bigger model. It is rare to see the memory layer isolated this cleanly from model scale, and it is the claim this article is built around.
I read the paper and cloned the repo to check it. The short version: the architecture is real and the mechanism is sensible, the headline numbers reproduce from the tables, and the benchmark story has two caveats — a judge mismatch and a set of self-reported competitor scores — that the paper is mostly honest about and the README oversells slightly.
Why memory past the context window is hard
The naive answer to "make the agent remember" is: paste the whole history into the prompt. This is the full-context baseline, and it fails in two ways at once. It does not fit — LongMemEval's larger setting is about 1.5 million tokens across roughly 500 sessions (reported) — and even when a long-context model swallows it, accuracy collapses. That is the 39.0% above: a capable 20B model, handed everything, still answers wrong most of the time because the one relevant sentence is diluted among hundreds of thousands of irrelevant ones.
The standard fix is RAG: chunk the history, embed it, and retrieve the top-k chunks for each query. This is better, and it is what most "agent memory" products are. The paper's argument — and it is a good one — is that top-k retrieval into a stateless model has three structural problems that no amount of embedding quality fixes:
- It blurs evidence and inference. A chunk that says "I think Postgres is the right call" and a chunk that says "we migrated to Postgres" are both just text in the store. The system cannot tell what the agent observed from what it believes, so it cannot reason about either reliably.
- It does not organize over long horizons. A flat pile of chunks has no notion that forty scattered mentions of one person are about that person, or that a belief stated in March was revised in June.
- It cannot explain itself. When the answer is "the top-k chunks happened to contain this", there is no trace of why the agent reasoned the way it did — which is exactly what a long-lived agent needs to be auditable.
Hindsight's bet is that memory should be a structured, first-class substrate the agent reasons over, not a thin retrieval shim around a model that stays stateless. The widget below makes the failure concrete: a fact stated in session 1, needed in session 20. Slide the context window and watch it fall off the edge; switch to Hindsight and watch it come back on its own merits.
“I don’t have any record of an allergy.” The fact was stated in session 1, which fell outside the last 6 sessions. It was truncated, so the model never saw it.
Full context loses the fact the moment the window stops reaching session 1 — and at real scale the window never reaches that far. Hindsight retained the fact as a structured memory when it was first said, so recall can find it nineteen sessions later regardless of how much was said in between. The rest of this piece is how that retain-and-recall actually works, and what "reflect" adds on top.
Four networks, three operations
The core abstraction is small. A memory bank (one "brain" per user, agent, or project) holds four logical networks, and three operations act on them.
The four networks each carry a distinct epistemic role — this is the "evidence vs. inference" separation made structural (reported, paper §3.1):
- World () — objective facts about the external world. "Alice works at Google in Mountain View on the AI team."
- Experience () — the agent's own first-person history. "I recommended Yosemite to Alice for hiking."
- Opinion () — subjective judgments, each a tuple of text, a confidence , and a timestamp. "Python is better for data science because of pandas" (confidence 0.85).
- Observation () — preference-neutral summaries of an entity, synthesized from world and experience facts. "Alice is a software engineer at Google specializing in machine learning."
World and experience are evidence. Opinions are belief, and they carry a confidence number so the agent can hold a view weakly. Observations are the agent's distilled profile of a thing. Keeping them in separate drawers is the design's first principle: a developer can see what the agent knows versus what it merely thinks.
The three operations are the public interface:
- Retain — ingest input , extract facts, classify each into one of the four networks, and update the memory graph.
- Recall — retrieve the most relevant facts for query within a token budget .
- Reflect — generate a response shaped by a behavioral profile , forming and updating opinions as it goes.
Retain and recall are implemented by a component the paper calls TEMPR (Temporal Entity Memory Priming Retrieval); reflect is CARA (Coherent Adaptive Reasoning Agents). The full data flow is one figure:

Retain: turning transcripts into a graph
Retain is the expensive, write-side operation, and it is where the structure gets built. An LLM reads each conversation and extracts 2–5 narrative facts per conversation (reported, paper §4.1.2) — coarse chunking on purpose. The paper contrasts this with sentence-level extraction: a narrative fact is self-contained and preserves cross-turn context ("over three messages, Alice decided to switch teams because of the AI work"), rather than shattering one decision into fragments that lose the why.
Each extracted fact becomes a memory unit — a tuple carrying a unique id, the bank, the narrative text, an embedding , an occurrence interval , a mention time , its network type, an optional confidence, and metadata. Entities are resolved to canonical form, and the facts are wired into a graph with four kinds of weighted edges: temporal, semantic, entity, and causal. That graph is what lets recall walk from a fact to its neighbors — same person, nearby time, caused-by — instead of only matching the query text.
In the background, retain also consolidates related facts into observations and runs opinion reinforcement (more on that below). None of this is free: a single retain is several LLM calls (extraction, entity work, belief updates), which is the real cost of the approach and the thing to watch if you run it at volume.
Recall: four retrievers, fused and reranked
Recall is the read side, and it is the part most worth copying. Instead of one similarity search, TEMPR runs four retrievers in parallel, each capturing a different notion of relevance (reported, paper §4.2.2):
- Semantic — cosine similarity over an HNSW
pgvectorindex. Catches paraphrase and conceptual match. - BM25 — lexical, full-text over a GIN index. Catches exact proper nouns and identifiers the embedding smears together. (If you want the ~35-year-old ranking function underneath this spelled out term by term, I wrote a whole piece on BM25.)
- Graph — spreading activation over the memory graph: start from the top semantic hits, propagate activation along edges with a decay factor, and let causal and entity edges carry more weight. This surfaces facts that do not look like the query at all but are connected to it through a shared entity or a causal chain.
- Temporal — when the query has a time constraint, a hybrid parser (rule-based date libraries, falling back to a small
google/flan-t5-smallsequence model for the awkward phrasings) resolves it to a date range and matches against occurrence intervals.
Four ranked lists come back, and the question is how to merge them when their scores are not comparable. The answer is Reciprocal Rank Fusion, which merges by rank rather than raw score:
where is fact 's rank in list (and a miss contributes nothing). The constant is the standard RRF default, and it is exactly what ships: I found reciprocal_rank_fusion(result_lists, k: int = 60) in hindsight-api-slim/hindsight_api/engine/search/fusion.py (measured, from the clone at commit 0be6c02). RRF's virtue is robustness — it does not need calibrated scores, it ignores missing items gracefully, and a fact that several lanes agree on rises naturally.
After fusion, a neural cross-encoder reranker (cross-encoder/ms-marco-MiniLM-L-6-v2 — also the shipped default, DEFAULT_RERANKER_LOCAL_MODEL in config.py, measured) jointly scores the query against each top candidate for a final ordering, and a greedy packing step fills the token budget . The widget above shows this lane by lane: three retrievers find the allergy fact, the time-range lane misses because the question carries no date, RRF sums what the hits contribute, and the reranker puts it first. (The ranks in the widget are illustrative; the RRF arithmetic uses the real .)
Reflect: disposition, and beliefs that move
Recall answers lookups. Reflect is for the questions that need the agent to form a view, and it is where Hindsight does something the other memory systems do not. Reflect takes a behavioral profile — three disposition traits, each on a 1-to-5 scale, plus a bias strength (reported, paper §5.2):
- Skepticism (1 = trusting, 5 = skeptical)
- Literalism (1 = flexible, 5 = literal)
- Empathy (1 = detached, 5 = empathetic)
- Bias strength — how hard the profile pushes, from fact-only at 0 to strongly opinionated at 1.
The same facts, run through two different profiles, produce two different opinions. The paper's example: given identical evidence about remote work, a trusting-flexible-empathetic agent concludes "remote work is a net positive… it removes commute time and creates space for flexible work", while a skeptical-literal-detached agent concludes "remote work risks undermining consistent performance… harder to maintain structure and oversight". Neither is a hallucination; the disposition decides which aspects of the evidence get weighted.

Opinions do not stay fixed. When new facts arrive, CARA finds related opinions (by shared entity or embedding similarity), classifies the new evidence as reinforce, weaken, contradict, or neutral, and nudges the confidence accordingly — up by a step for reinforcement, down by a step for weakening, down by twice that for a contradiction, unchanged for neutral. Small evidence moves a belief a little; repeated or contradicting evidence moves it a lot. This is the "20/20" in the title: the agent revises what it believed in light of what it later learns, and the revision is traceable because each opinion keeps its confidence and timestamp.
Do the numbers hold?
Yes, with footnotes. Here is the LongMemEval S-setting overall row (reported, paper Table 3; 500 questions, GPT-OSS-120B as judge for Hindsight's own rows):
| System | Backbone | Overall |
|---|---|---|
| Full-context | GPT-4o | 60.2 |
| Full-context | GPT-OSS-20B | 39.0 |
| Zep | GPT-4o | 71.2 |
| Supermemory | GPT-4o | 81.6 |
| Supermemory | GPT-5 | 84.6 |
| Supermemory | Gemini-3 | 85.2 |
| Hindsight | GPT-OSS-20B | 83.6 |
| Hindsight | GPT-OSS-120B | 89.0 |
| Hindsight | Gemini-3 | 91.4 |
The central claim checks: Hindsight on GPT-OSS-20B scores 83.6 against the same model's full-context 39.0, a gain of 44.6 points (reported as "+44.6", and , which is my own arithmetic — reasoned — and it matches). It also clears full-context GPT-4o at 60.2 with a far smaller model. The gains land exactly where the benchmark is hard: multi-session questions go from 21.1 to 79.7 and temporal reasoning from 31.6 to 79.7 (reported). Scaling the backbone takes it to 91.4 with Gemini-3 Pro, the best in the table.
On LoCoMo (reported, paper Table 4):
| Method | Single-Hop | Multi-Hop | Open Domain | Temporal | Overall |
|---|---|---|---|---|---|
| Backboard | 89.36 | 75.00 | 91.20 | 91.90 | 90.00 |
| Memobase (v0.0.37) | 70.92 | 46.88 | 77.17 | 85.05 | 75.78 |
| Zep | 74.11 | 66.04 | 67.71 | 79.79 | 75.14 |
| Mem0 | 67.13 | 51.15 | 72.93 | 55.51 | 66.88 |
| LangMem | 62.23 | 47.92 | 71.12 | 23.43 | 58.10 |
| Hindsight (OSS-20B) | 74.11 | 64.58 | 90.96 | 76.32 | 83.18 |
| Hindsight (OSS-120B) | 76.79 | 62.50 | 93.68 | 79.44 | 85.67 |
| Hindsight (Gemini-3) | 86.17 | 70.83 | 95.12 | 83.80 | 89.61 |
The abstract's "89.61% on LoCoMo vs. 75.78% for the strongest prior open system" is accurate once you know that 75.78% is Memobase, the best of the open systems (reported). Hindsight's 89.61 also takes the best Open Domain score (95.12). The one number above it — Backboard's 90.00 — comes with a flag the paper raises itself, and it is where the honesty discussion starts.
The two caveats behind the table
The competitor scores are reported, not reproduced. The LoCoMo baselines (Backboard, Memobase, Zep, Mem0, LangMem) are taken from Backboard's official leaderboard and are, in the paper's own words, "treated as reported reference points rather than our independently reproduced baselines" (reported, paper §7.3). The LongMemEval baselines come from the Supermemory technical report. So Backboard's 90.00 is a vendor's self-reported number that the authors could not reproduce — which is the right way to label it, and worth remembering before reading "state of the art" anywhere.
The judge is not the same across rows. Hindsight's own numbers are judged by GPT-OSS-120B; the borrowed LongMemEval baselines used that report's GPT-4o judge. Different judges score differently, so the Hindsight-vs-baseline gaps on LongMemEval are not strictly apples-to-apples. The paper is upfront that it standardized the judge only within its own runs.
That last point is a small lesson in reading benchmark pages. The paper's best LongMemEval score is 91.4; the README's live chart (as of January 2026) shows 94.6, with Supermemory at 85.92 — numbers that have moved past the paper on a continuously-updated leaderboard:

Neither number is wrong; they are answers to slightly different questions (a frozen paper run vs. a live leaderboard). But if you are going to quote one, quote the paper's, and say which backbone.
What it costs, and where it breaks
The approach is not free and the paper does not pretend otherwise. Retain is several LLM calls per conversation — extraction, entity resolution, opinion reinforcement — so write-heavy workloads pay for the structure up front, in exchange for cheaper, more accurate reads later. The token budgets used for recall in the experiments are, oddly, left as unfilled placeholders in the preprint (the setup section literally reads "set to … tokens"), so that one configuration detail is not reproducible from the paper alone. And as noted, the disposition system — the most interesting part — is switched to neutral for every reported number, so its benefit is argued by example, not measured.
None of that dents the central result. A structured, four-network memory with four parallel retrievers and a fusion-plus-rerank recall takes an open 20B model from "wrong most of the time" to "better than full-context GPT-4o" on long-horizon memory, with nothing changed but the memory. That the memory layer, not the model, carries the win is the claim — and on the evidence in the paper, it holds.
If you want the neighboring reads: Tencent's own open memory system, taken apart in the TencentDB agent-memory piece, reaches for the same RRF retrieval but sorts its tiers by cache stability rather than epistemics; VoiceMem is a voice-agent memory paper whose shipped adapter does not match its own abstract; the jev-semgrep piece argues the case against pure vector similarity that Hindsight's BM25 and graph lanes quietly concede; and if memory is one piece of the agent, Lilian Weng's harness framing and Prime Intellect's open agent are the loop it plugs into. A companion piece in this batch on context-language models takes up the other half of the question — what to do when the context window is the memory.