# Hindsight: +44 points on the same 20B backbone, from the memory layer

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/hindsight-agent-memory
> date: 2026-10-02
> tags: explainer, agents, agent-memory, retrieval, llm, benchmarks, open-source

A memory system trended to the top of GitHub, which does not happen often for a thing whose whole job is to remember what you said three sessions ago. Hindsight — from Vectorize — crossed 40,000 stars in days and is at about 44,420 as I write this (measured from the GitHub API, 2026-10-02). It is MIT-licensed Python, and it comes with a paper: *Hindsight is 20/20: Building Agent Memory that Retains, Recalls, and Reflects* (arXiv 2512.12818, submitted 2025-12-14).

<RepoCard repo="vectorize-io/hindsight" />

The star count is the news; the paper is the substance, and it contains one number that is worth the whole read. On LongMemEval, the same open-source 20B model scores **39.0%** with the full conversation pasted into its context, and **83.6%** when Hindsight sits in front of it (reported, paper Table 3). Same backbone, same judge, same questions — a 44.6-point swing that comes entirely from how the memory is organized, not from a bigger model. It is rare to see the memory layer isolated this cleanly from model scale, and it is the claim this article is built around.

I read the paper and cloned the repo to check it. The short version: the architecture is real and the mechanism is sensible, the headline numbers reproduce from the tables, and the benchmark story has two caveats — a judge mismatch and a set of self-reported competitor scores — that the paper is mostly honest about and the README oversells slightly.

## Why memory past the context window is hard

The naive answer to "make the agent remember" is: paste the whole history into the prompt. This is the **full-context** baseline, and it fails in two ways at once. It does not fit — LongMemEval's larger setting is about 1.5 million tokens across roughly 500 sessions (reported) — and even when a long-context model swallows it, accuracy collapses. That is the 39.0% above: a capable 20B model, handed everything, still answers wrong most of the time because the one relevant sentence is diluted among hundreds of thousands of irrelevant ones.

The standard fix is **RAG**: chunk the history, embed it, and retrieve the top-k chunks for each query. This is better, and it is what most "agent memory" products are. The paper's argument — and it is a good one — is that top-k retrieval into a stateless model has three structural problems that no amount of embedding quality fixes:

1. **It blurs evidence and inference.** A chunk that says "I think Postgres is the right call" and a chunk that says "we migrated to Postgres" are both just text in the store. The system cannot tell what the agent *observed* from what it *believes*, so it cannot reason about either reliably.
2. **It does not organize over long horizons.** A flat pile of chunks has no notion that forty scattered mentions of one person are *about* that person, or that a belief stated in March was revised in June.
3. **It cannot explain itself.** When the answer is "the top-k chunks happened to contain this", there is no trace of why the agent reasoned the way it did — which is exactly what a long-lived agent needs to be auditable.

Hindsight's bet is that memory should be a **structured, first-class substrate** the agent reasons over, not a thin retrieval shim around a model that stays stateless. The widget below makes the failure concrete: a fact stated in session 1, needed in session 20. Slide the context window and watch it fall off the edge; switch to Hindsight and watch it come back on its own merits.

<MemoryRecallDemo />

Full context loses the fact the moment the window stops reaching session 1 — and at real scale the window never reaches that far. Hindsight retained the fact as a structured memory when it was first said, so recall can find it nineteen sessions later regardless of how much was said in between. The rest of this piece is how that retain-and-recall actually works, and what "reflect" adds on top.

## Four networks, three operations

The core abstraction is small. A memory **bank** (one "brain" per user, agent, or project) holds four logical networks, and three operations act on them.

The four networks each carry a distinct epistemic role — this is the "evidence vs. inference" separation made structural (reported, paper §3.1):

- **World ($\mathcal{W}$)** — objective facts about the external world. *"Alice works at Google in Mountain View on the AI team."*
- **Experience ($\mathcal{B}$)** — the agent's own first-person history. *"I recommended Yosemite to Alice for hiking."*
- **Opinion ($\mathcal{O}$)** — subjective judgments, each a tuple $(t, c, \tau)$ of text, a confidence $c \in [0,1]$, and a timestamp. *"Python is better for data science because of pandas"* (confidence 0.85).
- **Observation ($\mathcal{S}$)** — preference-neutral summaries of an entity, synthesized from world and experience facts. *"Alice is a software engineer at Google specializing in machine learning."*

World and experience are evidence. Opinions are belief, and they carry a confidence number so the agent can hold a view *weakly*. Observations are the agent's distilled profile of a thing. Keeping them in separate drawers is the design's first principle: a developer can see what the agent knows versus what it merely thinks.

The three operations are the public interface:

- **Retain**$(B, D) \rightarrow \mathcal{M}'$ — ingest input $D$, extract facts, classify each into one of the four networks, and update the memory graph.
- **Recall**$(B, Q, k) \rightarrow \{f_1, \dots, f_n\}$ — retrieve the most relevant facts for query $Q$ within a token budget $k$.
- **Reflect**$(B, Q, \Theta) \rightarrow (r, \mathcal{O}')$ — generate a response shaped by a behavioral profile $\Theta$, forming and updating opinions as it goes.

Retain and recall are implemented by a component the paper calls **TEMPR** (Temporal Entity Memory Priming Retrieval); reflect is **CARA** (Coherent Adaptive Reasoning Agents). The full data flow is one figure:

<Figure src="https://ai.thesatyajit.com/articles/hindsight-agent-memory/fig1.png" alt="End-to-end Hindsight architecture: input (token budget k, query Q, corpus D) feeds TEMPR's retain pipeline (fact extraction, embedding, entity resolution, link construction with temporal/semantic/entity/causal links) into a memory bank B holding the four networks W, B, O, S and a memory graph; TEMPR's recall pipeline runs four-way parallel retrieval (semantic, BM25, graph, temporal), RRF fusion, and cross-encoder reranking; CARA's reflect takes the recalled facts plus an agent profile with disposition parameters to generate a response r and update opinions O prime." caption="The whole system on one page: retain builds the four-network graph, recall runs four retrievers into RRF and a reranker, reflect conditions generation on a disposition profile and writes opinions back (Hindsight, Figure 2)." />

## Retain: turning transcripts into a graph

Retain is the expensive, write-side operation, and it is where the structure gets built. An LLM reads each conversation and extracts **2–5 narrative facts** per conversation (reported, paper §4.1.2) — coarse chunking on purpose. The paper contrasts this with sentence-level extraction: a narrative fact is self-contained and preserves cross-turn context ("over three messages, Alice decided to switch teams because of the AI work"), rather than shattering one decision into fragments that lose the *why*.

Each extracted fact becomes a memory unit — a tuple carrying a unique id, the bank, the narrative text, an embedding $v \in \mathbb{R}^d$, an occurrence interval $(\tau_s, \tau_e)$, a mention time $\tau_m$, its network type, an optional confidence, and metadata. Entities are resolved to canonical form, and the facts are wired into a graph $\mathcal{G} = (V, E)$ with four kinds of weighted edges: **temporal, semantic, entity, and causal**. That graph is what lets recall walk from a fact to its neighbors — same person, nearby time, caused-by — instead of only matching the query text.

In the background, retain also consolidates related facts into observations and runs **opinion reinforcement** (more on that below). None of this is free: a single retain is several LLM calls (extraction, entity work, belief updates), which is the real cost of the approach and the thing to watch if you run it at volume.

## Recall: four retrievers, fused and reranked

Recall is the read side, and it is the part most worth copying. Instead of one similarity search, TEMPR runs **four retrievers in parallel**, each capturing a different notion of relevance (reported, paper §4.2.2):

- **Semantic** — cosine similarity over an HNSW `pgvector` index. Catches paraphrase and conceptual match.
- **BM25** — lexical, full-text over a GIN index. Catches exact proper nouns and identifiers the embedding smears together. (If you want the ~35-year-old ranking function underneath this spelled out term by term, I wrote a whole piece on [BM25](/articles/bm25).)
- **Graph** — spreading activation over the memory graph: start from the top semantic hits, propagate activation along edges with a decay factor, and let causal and entity edges carry more weight. This surfaces facts that do not look like the query at all but are connected to it through a shared entity or a causal chain.
- **Temporal** — when the query has a time constraint, a hybrid parser (rule-based date libraries, falling back to a small `google/flan-t5-small` sequence model for the awkward phrasings) resolves it to a date range and matches against occurrence intervals.

Four ranked lists come back, and the question is how to merge them when their scores are not comparable. The answer is **Reciprocal Rank Fusion**, which merges by *rank* rather than raw score:

$$\mathrm{RRF}(f) = \sum_{i=1}^{4} \frac{1}{k + r_i(f)}, \qquad k = 60$$

where $r_i(f)$ is fact $f$'s rank in list $i$ (and a miss contributes nothing). The constant $k = 60$ is the standard RRF default, and it is exactly what ships: I found `reciprocal_rank_fusion(result_lists, k: int = 60)` in `hindsight-api-slim/hindsight_api/engine/search/fusion.py` (measured, from the clone at commit `0be6c02`). RRF's virtue is robustness — it does not need calibrated scores, it ignores missing items gracefully, and a fact that several lanes agree on rises naturally.

After fusion, a neural **cross-encoder reranker** (`cross-encoder/ms-marco-MiniLM-L-6-v2` — also the shipped default, `DEFAULT_RERANKER_LOCAL_MODEL` in `config.py`, measured) jointly scores the query against each top candidate for a final ordering, and a greedy packing step fills the token budget $k$. The widget above shows this lane by lane: three retrievers find the allergy fact, the time-range lane misses because the question carries no date, RRF sums what the hits contribute, and the reranker puts it first. (The ranks in the widget are illustrative; the RRF arithmetic uses the real $k = 60$.)

## Reflect: disposition, and beliefs that move

Recall answers lookups. **Reflect** is for the questions that need the agent to form a view, and it is where Hindsight does something the other memory systems do not. Reflect takes a behavioral profile $\Theta = (S, L, E, \beta)$ — three **disposition traits**, each on a 1-to-5 scale, plus a bias strength (reported, paper §5.2):

- **Skepticism** $S$ (1 = trusting, 5 = skeptical)
- **Literalism** $L$ (1 = flexible, 5 = literal)
- **Empathy** $E$ (1 = detached, 5 = empathetic)
- **Bias strength** $\beta \in [0,1]$ — how hard the profile pushes, from fact-only at 0 to strongly opinionated at 1.

The same facts, run through two different profiles, produce two different opinions. The paper's example: given identical evidence about remote work, a trusting-flexible-empathetic agent concludes *"remote work is a net positive… it removes commute time and creates space for flexible work"*, while a skeptical-literal-detached agent concludes *"remote work risks undermining consistent performance… harder to maintain structure and oversight"*. Neither is a hallucination; the disposition decides which aspects of the evidence get weighted.

<Figure src="https://ai.thesatyajit.com/articles/hindsight-agent-memory/fig2.png" alt="CARA's reflect loop: an input query triggers recall of memories, which builds context; the agent loads its profile (disposition traits skepticism, literalism, empathy, plus background); LLM generation is conditioned on that disposition; the generated response becomes the final response and also drives create-or-update of opinion and observation memories, each with an adjust-confidence step writing back to the memory store." caption="Reflect recalls, loads the bank's disposition and background, generates a response conditioned on them, and writes new or updated opinions and observations back with adjusted confidence (Hindsight, Figure 4)." />

Opinions do not stay fixed. When new facts arrive, CARA finds related opinions (by shared entity or embedding similarity), classifies the new evidence as *reinforce*, *weaken*, *contradict*, or *neutral*, and nudges the confidence accordingly — up by a step for reinforcement, down by a step for weakening, down by twice that for a contradiction, unchanged for neutral. Small evidence moves a belief a little; repeated or contradicting evidence moves it a lot. This is the "20/20" in the title: the agent revises what it believed in light of what it later learns, and the revision is traceable because each opinion keeps its confidence and timestamp.

<Callout type="note">
For the benchmarks, the disposition machinery is essentially turned off: the paper runs with neutral traits (skepticism, literalism, empathy all set to 3) and low bias strength (0.2), because LongMemEval and LoCoMo test factual recall, not viewpoint consistency (reported, paper §7.3). So the headline accuracy numbers are a test of *retain and recall*, not of reflect. The disposition system is the paper's most novel idea and its least measured one.
</Callout>

## Do the numbers hold?

Yes, with footnotes. Here is the LongMemEval S-setting overall row (reported, paper Table 3; 500 questions, GPT-OSS-120B as judge for Hindsight's own rows):

| System | Backbone | Overall |
|---|---|---|
| Full-context | GPT-4o | 60.2 |
| Full-context | GPT-OSS-20B | 39.0 |
| Zep | GPT-4o | 71.2 |
| Supermemory | GPT-4o | 81.6 |
| Supermemory | GPT-5 | 84.6 |
| Supermemory | Gemini-3 | 85.2 |
| **Hindsight** | **GPT-OSS-20B** | **83.6** |
| **Hindsight** | GPT-OSS-120B | 89.0 |
| **Hindsight** | Gemini-3 | **91.4** |

The central claim checks: Hindsight on GPT-OSS-20B scores 83.6 against the same model's full-context 39.0, a gain of 44.6 points (reported as "+44.6", and $83.6 - 39.0 = 44.6$, which is my own arithmetic — reasoned — and it matches). It also clears full-context GPT-4o at 60.2 with a far smaller model. The gains land exactly where the benchmark is hard: multi-session questions go from 21.1 to 79.7 and temporal reasoning from 31.6 to 79.7 (reported). Scaling the backbone takes it to 91.4 with Gemini-3 Pro, the best in the table.

On LoCoMo (reported, paper Table 4):

| Method | Single-Hop | Multi-Hop | Open Domain | Temporal | Overall |
|---|---|---|---|---|---|
| Backboard | 89.36 | 75.00 | 91.20 | 91.90 | 90.00 |
| Memobase (v0.0.37) | 70.92 | 46.88 | 77.17 | 85.05 | 75.78 |
| Zep | 74.11 | 66.04 | 67.71 | 79.79 | 75.14 |
| Mem0 | 67.13 | 51.15 | 72.93 | 55.51 | 66.88 |
| LangMem | 62.23 | 47.92 | 71.12 | 23.43 | 58.10 |
| **Hindsight** (OSS-20B) | 74.11 | 64.58 | 90.96 | 76.32 | 83.18 |
| **Hindsight** (OSS-120B) | 76.79 | 62.50 | 93.68 | 79.44 | 85.67 |
| **Hindsight** (Gemini-3) | 86.17 | 70.83 | 95.12 | 83.80 | **89.61** |

The abstract's "89.61% on LoCoMo vs. 75.78% for the strongest prior open system" is accurate once you know that 75.78% is **Memobase**, the best of the *open* systems (reported). Hindsight's 89.61 also takes the best Open Domain score (95.12). The one number above it — Backboard's 90.00 — comes with a flag the paper raises itself, and it is where the honesty discussion starts.

## The two caveats behind the table

**The competitor scores are reported, not reproduced.** The LoCoMo baselines (Backboard, Memobase, Zep, Mem0, LangMem) are taken from Backboard's official leaderboard and are, in the paper's own words, "treated as reported reference points rather than our independently reproduced baselines" (reported, paper §7.3). The LongMemEval baselines come from the Supermemory technical report. So Backboard's 90.00 is a vendor's self-reported number that the authors could not reproduce — which is the right way to label it, and worth remembering before reading "state of the art" anywhere.

**The judge is not the same across rows.** Hindsight's own numbers are judged by GPT-OSS-120B; the borrowed LongMemEval baselines used that report's GPT-4o judge. Different judges score differently, so the Hindsight-vs-baseline gaps on LongMemEval are not strictly apples-to-apples. The paper is upfront that it standardized the judge only *within* its own runs.

<Callout type="warning">
The repo's README claims the benchmark data was "independently reproduced by research collaborators at the Virginia Tech Sanghani Center… and The Washington Post", and that other scores are self-reported by vendors. The self-reported part is in the paper; the third-party reproduction is a README claim with no linked artifact I could open, so I cannot verify it — treat it as a claim, not a confirmed result. Separately, the README's live benchmark chart reports a *higher* Hindsight number than the paper does.
</Callout>

That last point is a small lesson in reading benchmark pages. The paper's best LongMemEval score is 91.4; the README's live chart (as of January 2026) shows 94.6, with Supermemory at 85.92 — numbers that have moved past the paper on a continuously-updated leaderboard:

<Figure src="https://ai.thesatyajit.com/articles/hindsight-agent-memory/fig3.png" alt="Horizontal bar chart of LongMemEval overall scores from the Hindsight README: GPT-4o 60.2%, Zep 71.2%, SuperMemory 85.92%, Hindsight 94.6%, labelled 'LongMemEval Overall Score 94.6%'." caption="The project's live README chart, which reports Hindsight at 94.6% — higher than the paper's 91.4%, because the leaderboard keeps updating. Competitor scores here are vendor self-reported (Hindsight, project README)." />

Neither number is wrong; they are answers to slightly different questions (a frozen paper run vs. a live leaderboard). But if you are going to quote one, quote the paper's, and say which backbone.

## What it costs, and where it breaks

The approach is not free and the paper does not pretend otherwise. Retain is several LLM calls per conversation — extraction, entity resolution, opinion reinforcement — so write-heavy workloads pay for the structure up front, in exchange for cheaper, more accurate reads later. The token budgets used for recall in the experiments are, oddly, left as unfilled placeholders in the preprint (the setup section literally reads "set to … tokens"), so that one configuration detail is not reproducible from the paper alone. And as noted, the disposition system — the most interesting part — is switched to neutral for every reported number, so its benefit is argued by example, not measured.

None of that dents the central result. A structured, four-network memory with four parallel retrievers and a fusion-plus-rerank recall takes an open 20B model from "wrong most of the time" to "better than full-context GPT-4o" on long-horizon memory, with nothing changed but the memory. That the memory layer, not the model, carries the win is the claim — and on the evidence in the paper, it holds.

If you want the neighboring reads: Tencent's own open memory system, taken apart in [the TencentDB agent-memory piece](/articles/tencentdb-agent-memory), reaches for the same RRF retrieval but sorts its tiers by cache stability rather than epistemics; [VoiceMem](/articles/voicemem) is a voice-agent memory paper whose shipped adapter does not match its own abstract; [the jev-semgrep piece](/articles/search-by-meaning) argues the case *against* pure vector similarity that Hindsight's BM25 and graph lanes quietly concede; and if memory is one piece of the agent, [Lilian Weng's harness framing](/articles/agent-harness) and [Prime Intellect's open agent](/articles/prime-agent) are the loop it plugs into. A companion piece in this batch on context-language models takes up the other half of the question — what to do when the context window is the memory.
