~/satyajit

Hindsight: +44 points on the same 20B backbone, from the memory layer

mdjsonmcp

2026-10-02 · 16 min · explainer · agents · agent-memory · retrieval · llm · benchmarks · open-source

A memory system trended to the top of GitHub, which does not happen often for a thing whose whole job is to remember what you said three sessions ago. Hindsight — from Vectorize — crossed 40,000 stars in days and is at about 44,420 as I write this (measured from the GitHub API, 2026-10-02). It is MIT-licensed Python, and it comes with a paper: Hindsight is 20/20: Building Agent Memory that Retains, Recalls, and Reflects (arXiv 2512.12818, submitted 2025-12-14).

vectorize-io/hindsight@0be6c02 · snapshot 2026-10-02
tracked files
4,965
license
MIT
branch
main
tests
1206 files
source
32.8 MB
commit date
2026-10-01
source by language
Python23.3 MB(1929)TypeScript5.8 MB(732)Go2.4 MB(245)Rust565.2 kB(36)Shell327.6 kB(73)JavaScript263.9 kB(68)CSS181.5 kB(25)

by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded

local clone, 2026-10-02 at 0be6c02 — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow, testFileCount

shallow clone: counts describe the pinned tree, not the history

The star count is the news; the paper is the substance, and it contains one number that is worth the whole read. On LongMemEval, the same open-source 20B model scores 39.0% with the full conversation pasted into its context, and 83.6% when Hindsight sits in front of it (reported, paper Table 3). Same backbone, same judge, same questions — a 44.6-point swing that comes entirely from how the memory is organized, not from a bigger model. It is rare to see the memory layer isolated this cleanly from model scale, and it is the claim this article is built around.

I read the paper and cloned the repo to check it. The short version: the architecture is real and the mechanism is sensible, the headline numbers reproduce from the tables, and the benchmark story has two caveats — a judge mismatch and a set of self-reported competitor scores — that the paper is mostly honest about and the README oversells slightly.

Why memory past the context window is hard

The naive answer to "make the agent remember" is: paste the whole history into the prompt. This is the full-context baseline, and it fails in two ways at once. It does not fit — LongMemEval's larger setting is about 1.5 million tokens across roughly 500 sessions (reported) — and even when a long-context model swallows it, accuracy collapses. That is the 39.0% above: a capable 20B model, handed everything, still answers wrong most of the time because the one relevant sentence is diluted among hundreds of thousands of irrelevant ones.

The standard fix is RAG: chunk the history, embed it, and retrieve the top-k chunks for each query. This is better, and it is what most "agent memory" products are. The paper's argument — and it is a good one — is that top-k retrieval into a stateless model has three structural problems that no amount of embedding quality fixes:

  1. It blurs evidence and inference. A chunk that says "I think Postgres is the right call" and a chunk that says "we migrated to Postgres" are both just text in the store. The system cannot tell what the agent observed from what it believes, so it cannot reason about either reliably.
  2. It does not organize over long horizons. A flat pile of chunks has no notion that forty scattered mentions of one person are about that person, or that a belief stated in March was revised in June.
  3. It cannot explain itself. When the answer is "the top-k chunks happened to contain this", there is no trace of why the agent reasoned the way it did — which is exactly what a long-lived agent needs to be auditable.

Hindsight's bet is that memory should be a structured, first-class substrate the agent reasons over, not a thin retrieval shim around a model that stays stateless. The widget below makes the failure concrete: a fact stated in session 1, needed in session 20. Slide the context window and watch it fall off the edge; switch to Hindsight and watch it come back on its own merits.

session 1 says it; session 20 needs it
1234567891011121314151617181920
1 “Mira is allergic to penicillin.”20 “which antibiotic should we avoid?”
context window: last 6 sessions
raw transcript 40,000 tok · kept 12,000 tok · budget 16,000 tok
answer (wrong)

“I don’t have any record of an allergy.” The fact was stated in session 1, which fell outside the last 6 sessions. It was truncated, so the model never saw it.

Full context loses the fact the moment the window stops reaching session 1 — and at real scale the window never reaches that far. Hindsight retained the fact as a structured memory when it was first said, so recall can find it nineteen sessions later regardless of how much was said in between. The rest of this piece is how that retain-and-recall actually works, and what "reflect" adds on top.

Four networks, three operations

The core abstraction is small. A memory bank (one "brain" per user, agent, or project) holds four logical networks, and three operations act on them.

The four networks each carry a distinct epistemic role — this is the "evidence vs. inference" separation made structural (reported, paper §3.1):

World and experience are evidence. Opinions are belief, and they carry a confidence number so the agent can hold a view weakly. Observations are the agent's distilled profile of a thing. Keeping them in separate drawers is the design's first principle: a developer can see what the agent knows versus what it merely thinks.

The three operations are the public interface:

Retain and recall are implemented by a component the paper calls TEMPR (Temporal Entity Memory Priming Retrieval); reflect is CARA (Coherent Adaptive Reasoning Agents). The full data flow is one figure:

End-to-end Hindsight architecture: input (token budget k, query Q, corpus D) feeds TEMPR's retain pipeline (fact extraction, embedding, entity resolution, link construction with temporal/semantic/entity/causal links) into a memory bank B holding the four networks W, B, O, S and a memory graph; TEMPR's recall pipeline runs four-way parallel retrieval (semantic, BM25, graph, temporal), RRF fusion, and cross-encoder reranking; CARA's reflect takes the recalled facts plus an agent profile with disposition parameters to generate a response r and update opinions O prime.
The whole system on one page: retain builds the four-network graph, recall runs four retrievers into RRF and a reranker, reflect conditions generation on a disposition profile and writes opinions back (Hindsight, Figure 2).

Retain: turning transcripts into a graph

Retain is the expensive, write-side operation, and it is where the structure gets built. An LLM reads each conversation and extracts 2–5 narrative facts per conversation (reported, paper §4.1.2) — coarse chunking on purpose. The paper contrasts this with sentence-level extraction: a narrative fact is self-contained and preserves cross-turn context ("over three messages, Alice decided to switch teams because of the AI work"), rather than shattering one decision into fragments that lose the why.

Each extracted fact becomes a memory unit — a tuple carrying a unique id, the bank, the narrative text, an embedding v∈Rdv \in \mathbb{R}^d, an occurrence interval (τs,τe)(\tau_s, \tau_e), a mention time τm\tau_m, its network type, an optional confidence, and metadata. Entities are resolved to canonical form, and the facts are wired into a graph G=(V,E)\mathcal{G} = (V, E) with four kinds of weighted edges: temporal, semantic, entity, and causal. That graph is what lets recall walk from a fact to its neighbors — same person, nearby time, caused-by — instead of only matching the query text.

In the background, retain also consolidates related facts into observations and runs opinion reinforcement (more on that below). None of this is free: a single retain is several LLM calls (extraction, entity work, belief updates), which is the real cost of the approach and the thing to watch if you run it at volume.

Recall: four retrievers, fused and reranked

Recall is the read side, and it is the part most worth copying. Instead of one similarity search, TEMPR runs four retrievers in parallel, each capturing a different notion of relevance (reported, paper §4.2.2):

Four ranked lists come back, and the question is how to merge them when their scores are not comparable. The answer is Reciprocal Rank Fusion, which merges by rank rather than raw score:

RRF(f)=∑i=141k+ri(f),k=60\mathrm{RRF}(f) = \sum_{i=1}^{4} \frac{1}{k + r_i(f)}, \qquad k = 60

where ri(f)r_i(f) is fact ff's rank in list ii (and a miss contributes nothing). The constant k=60k = 60 is the standard RRF default, and it is exactly what ships: I found reciprocal_rank_fusion(result_lists, k: int = 60) in hindsight-api-slim/hindsight_api/engine/search/fusion.py (measured, from the clone at commit 0be6c02). RRF's virtue is robustness — it does not need calibrated scores, it ignores missing items gracefully, and a fact that several lanes agree on rises naturally.

After fusion, a neural cross-encoder reranker (cross-encoder/ms-marco-MiniLM-L-6-v2 — also the shipped default, DEFAULT_RERANKER_LOCAL_MODEL in config.py, measured) jointly scores the query against each top candidate for a final ordering, and a greedy packing step fills the token budget kk. The widget above shows this lane by lane: three retrievers find the allergy fact, the time-range lane misses because the question carries no date, RRF sums what the hits contribute, and the reranker puts it first. (The ranks in the widget are illustrative; the RRF arithmetic uses the real k=60k = 60.)

Reflect: disposition, and beliefs that move

Recall answers lookups. Reflect is for the questions that need the agent to form a view, and it is where Hindsight does something the other memory systems do not. Reflect takes a behavioral profile Θ=(S,L,E,β)\Theta = (S, L, E, \beta) — three disposition traits, each on a 1-to-5 scale, plus a bias strength (reported, paper §5.2):

The same facts, run through two different profiles, produce two different opinions. The paper's example: given identical evidence about remote work, a trusting-flexible-empathetic agent concludes "remote work is a net positive… it removes commute time and creates space for flexible work", while a skeptical-literal-detached agent concludes "remote work risks undermining consistent performance… harder to maintain structure and oversight". Neither is a hallucination; the disposition decides which aspects of the evidence get weighted.

CARA's reflect loop: an input query triggers recall of memories, which builds context; the agent loads its profile (disposition traits skepticism, literalism, empathy, plus background); LLM generation is conditioned on that disposition; the generated response becomes the final response and also drives create-or-update of opinion and observation memories, each with an adjust-confidence step writing back to the memory store.
Reflect recalls, loads the bank's disposition and background, generates a response conditioned on them, and writes new or updated opinions and observations back with adjusted confidence (Hindsight, Figure 4).

Opinions do not stay fixed. When new facts arrive, CARA finds related opinions (by shared entity or embedding similarity), classifies the new evidence as reinforce, weaken, contradict, or neutral, and nudges the confidence accordingly — up by a step for reinforcement, down by a step for weakening, down by twice that for a contradiction, unchanged for neutral. Small evidence moves a belief a little; repeated or contradicting evidence moves it a lot. This is the "20/20" in the title: the agent revises what it believed in light of what it later learns, and the revision is traceable because each opinion keeps its confidence and timestamp.

Do the numbers hold?

Yes, with footnotes. Here is the LongMemEval S-setting overall row (reported, paper Table 3; 500 questions, GPT-OSS-120B as judge for Hindsight's own rows):

SystemBackboneOverall
Full-contextGPT-4o60.2
Full-contextGPT-OSS-20B39.0
ZepGPT-4o71.2
SupermemoryGPT-4o81.6
SupermemoryGPT-584.6
SupermemoryGemini-385.2
HindsightGPT-OSS-20B83.6
HindsightGPT-OSS-120B89.0
HindsightGemini-391.4

The central claim checks: Hindsight on GPT-OSS-20B scores 83.6 against the same model's full-context 39.0, a gain of 44.6 points (reported as "+44.6", and 83.6−39.0=44.683.6 - 39.0 = 44.6, which is my own arithmetic — reasoned — and it matches). It also clears full-context GPT-4o at 60.2 with a far smaller model. The gains land exactly where the benchmark is hard: multi-session questions go from 21.1 to 79.7 and temporal reasoning from 31.6 to 79.7 (reported). Scaling the backbone takes it to 91.4 with Gemini-3 Pro, the best in the table.

On LoCoMo (reported, paper Table 4):

MethodSingle-HopMulti-HopOpen DomainTemporalOverall
Backboard89.3675.0091.2091.9090.00
Memobase (v0.0.37)70.9246.8877.1785.0575.78
Zep74.1166.0467.7179.7975.14
Mem067.1351.1572.9355.5166.88
LangMem62.2347.9271.1223.4358.10
Hindsight (OSS-20B)74.1164.5890.9676.3283.18
Hindsight (OSS-120B)76.7962.5093.6879.4485.67
Hindsight (Gemini-3)86.1770.8395.1283.8089.61

The abstract's "89.61% on LoCoMo vs. 75.78% for the strongest prior open system" is accurate once you know that 75.78% is Memobase, the best of the open systems (reported). Hindsight's 89.61 also takes the best Open Domain score (95.12). The one number above it — Backboard's 90.00 — comes with a flag the paper raises itself, and it is where the honesty discussion starts.

The two caveats behind the table

The competitor scores are reported, not reproduced. The LoCoMo baselines (Backboard, Memobase, Zep, Mem0, LangMem) are taken from Backboard's official leaderboard and are, in the paper's own words, "treated as reported reference points rather than our independently reproduced baselines" (reported, paper §7.3). The LongMemEval baselines come from the Supermemory technical report. So Backboard's 90.00 is a vendor's self-reported number that the authors could not reproduce — which is the right way to label it, and worth remembering before reading "state of the art" anywhere.

The judge is not the same across rows. Hindsight's own numbers are judged by GPT-OSS-120B; the borrowed LongMemEval baselines used that report's GPT-4o judge. Different judges score differently, so the Hindsight-vs-baseline gaps on LongMemEval are not strictly apples-to-apples. The paper is upfront that it standardized the judge only within its own runs.

That last point is a small lesson in reading benchmark pages. The paper's best LongMemEval score is 91.4; the README's live chart (as of January 2026) shows 94.6, with Supermemory at 85.92 — numbers that have moved past the paper on a continuously-updated leaderboard:

Horizontal bar chart of LongMemEval overall scores from the Hindsight README: GPT-4o 60.2%, Zep 71.2%, SuperMemory 85.92%, Hindsight 94.6%, labelled 'LongMemEval Overall Score 94.6%'.
The project's live README chart, which reports Hindsight at 94.6% — higher than the paper's 91.4%, because the leaderboard keeps updating. Competitor scores here are vendor self-reported (Hindsight, project README).

Neither number is wrong; they are answers to slightly different questions (a frozen paper run vs. a live leaderboard). But if you are going to quote one, quote the paper's, and say which backbone.

What it costs, and where it breaks

The approach is not free and the paper does not pretend otherwise. Retain is several LLM calls per conversation — extraction, entity resolution, opinion reinforcement — so write-heavy workloads pay for the structure up front, in exchange for cheaper, more accurate reads later. The token budgets used for recall in the experiments are, oddly, left as unfilled placeholders in the preprint (the setup section literally reads "set to … tokens"), so that one configuration detail is not reproducible from the paper alone. And as noted, the disposition system — the most interesting part — is switched to neutral for every reported number, so its benefit is argued by example, not measured.

None of that dents the central result. A structured, four-network memory with four parallel retrievers and a fusion-plus-rerank recall takes an open 20B model from "wrong most of the time" to "better than full-context GPT-4o" on long-horizon memory, with nothing changed but the memory. That the memory layer, not the model, carries the win is the claim — and on the evidence in the paper, it holds.

If you want the neighboring reads: Tencent's own open memory system, taken apart in the TencentDB agent-memory piece, reaches for the same RRF retrieval but sorts its tiers by cache stability rather than epistemics; VoiceMem is a voice-agent memory paper whose shipped adapter does not match its own abstract; the jev-semgrep piece argues the case against pure vector similarity that Hindsight's BM25 and graph lanes quietly concede; and if memory is one piece of the agent, Lilian Weng's harness framing and Prime Intellect's open agent are the loop it plugs into. A companion piece in this batch on context-language models takes up the other half of the question — what to do when the context window is the memory.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "Hindsight: +44 points on the same 20B backbone, from the memory layer", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026hindsightagentmemory,
  author = {Satyajit Ghana},
  title  = {Hindsight: +44 points on the same 20B backbone, from the memory layer},
  url    = {https://ai.thesatyajit.com/articles/hindsight-agent-memory},
  year   = {2026}
}
share