2026-10-06 · 20 min · agents · agent-memory · benchmarks
Why read this
Notabletop 60%Why always-on failure lessons hurt and how Sentry's detector, label-keyed retrieval and verified writes fix it, with every headline average re-derived.
- Original analysis
- A lasting reference
- A new technique
Agents & harnessesAPI onlyMITResearch paper
How this was scored
- Is it new?
- 2 of 3: A real new idea, method or capability
- Can I trust it?
- 2 of 3: Measures key facts from files, code or configs
- Can I run it?
- 0 of 3: Closed, nothing to run
- Will I understand it?
- 2 of 3: Mechanism from first principles with figures
- Can I act on it?
- 2 of 3: A concrete recipe, numbers or comparison
- Will it last?
- 2 of 3: A reference for a year or more
- Does it affect many?
- 1 of 3: A specialist community
- Only here?
- 2 of 3: A teardown or measurement few others did
Score 59 of 100, ranked 245 of 445 rated articles. Each question is answered 0–3 by hand, and a 3 is rare. How articles are scored
The usual way to make an agent learn from its mistakes is to write them down and put the notes in its prompt. ACE, Stanford's Agentic Context Engineering, keeps an evolving playbook of itemised advice that a Generator, Reflector and Curator edit after every task.
The ACE team's new paper, Sentry (arXiv 2610.02994, Changxiu Ji, Amy Lu, Qizheng Zhang and Kunle Olukotun), starts from an awkward finding about their own method. The failure lessons in that playbook make the agent worse on tasks where the failure never happens. Take the failure lessons out of ACE's playbook and held-out WebShop reward goes up, from 0.335 to 0.391 (reported).
The paper's conclusion is not "stop storing lessons". It is that a lesson of the form if X goes wrong, do Y is conditional knowledge, and it should reach the agent only when X is actually going wrong. Sentry is the machinery that makes that happen: a failure detector, a playbook kept outside the agent's context, retrieval keyed by failure label, and a write rule that admits a lesson only after the repair is judged to have worked.
Two ways failure knowledge goes wrong

Context evolution conditions on the task. ACE, Reflexion, AWM, ExpeL and ReasoningBank all choose their lessons before execution starts. Even the retrieval-based ones pick a subset at task start and leave it in context for the whole trajectory. But whether "broaden the query after repeated empty searches" applies depends on the agent's state at one particular step, which nobody knows at task start. So the agent has to decide at every step which of its rules apply, and it gets that wrong.
The paper's example, from its appendix trajectory for WebShop task 4494: an earlier task (6594) looped on search[women's pumps] twenty times for zero reward. ACE distilled three rules from it: no viable result, broaden the query; no match on page 1, check the next pages; results still irrelevant, refine the query. On task 4494 the agent searches for 20-inch hair extensions under $80, and the first result is a 20-inch item at $79.63. The agent writes "Item 1 fits the description perfectly. Let me also go through item 2, item 3, item 4 to make sure there is nothing better…", emits no action, and scores 0. The rules were written for a search that fails. This search worked, and the rules still had the agent's attention.
Runtime intervention conditions on the state, but learns nothing. AgentGuard, AgentForesight, Wink and AgentFixer watch the trajectory and inject a warning when something looks wrong. That is the right conditioning variable. Their knowledge is fixed, though. The paper's AgentGuard implementation warns when the invalid-transition rate exceeds 0.25, the no-progress rate exceeds 0.75, or the same no-progress action repeats three times in a row. On task 3980 its "No progress detected. Take a state-changing action." breaks a real search loop and the agent buys the right item. On task 4585 the same warning fires after the agent clicks a colour and a size, because selecting options leaves the product page unchanged. The agent obeys, clicks < Prev, loses both options, and buys the item without them for a reward of 0.6. Nothing in the rule can learn that "page unchanged" meant progress this time.
- • no viable result → broaden the query
- • no match on page 1 → check the next pages
- • results still irrelevant → refine the query
- search[hair extensions 20 inch]→ 1. QAZPL Deep Curly Virgin Human Hair, 20 inch, $79.63 · 2. … 3. … 4. …monitor: none: ACE has no runtime monitor
- think: “Item 1 fits the description perfectly. Let me also go through item 2, item 3, item 4 to make sure there is nothing better…”→ no action emittedmonitor: none
- (episode ends)→ reward 0 · the correct click[B08X4YD5WH] was never takenmonitor: none
The ACE and AgentGuard episodes are condensed from the paper's Figure 4. The Sentry episode is my illustration, since the paper prints no full Sentry trajectory: the ACE task replayed under Sentry's published rules, with two lesson texts verbatim from the paper's example playbook and no reward. Watch the dots. ACE has three recovery rules in the prompt at every step. Sentry has none until its detector diagnoses over-exploration, and then two.
How much the always-on lessons cost
The paper measures the exposure cost twice, both on held-out tasks (memory initialised on 50 tasks, scored on 50 disjoint ones, five runs).
| Experiment | WebShop reward | Mind2Web Replay completion |
|---|---|---|
| ACE, full playbook | 0.335 | 0.562 |
| ACE, failure lessons pruned | 0.391 | 0.643 |
| Sentry, frozen playbook, failure-triggered only | 0.364 | 0.651 |
| same, plus the full playbook in persistent context | 0.296 | 0.599 |
All eight values are reported. Pruning helps ACE by 17% and 14% relative (reasoned: 0.391 / 0.335 and 0.643 / 0.562). The second comparison is the cleaner one. The two Sentry variants use the same frozen playbook, the same detector, the same retrieval and the same verifier, and failure-triggered retrieval stays on in both. The only difference is whether the whole playbook also sits in the agent's background context. Adding it costs 0.068 reward on WebShop and 0.052 completion on Mind2Web Replay (reported), a 19% relative drop on WebShop (reasoned). The relevant lessons are still available on demand in both arms. Showing the rest of them all the time is what costs.
One nuance the paper's framing skips: the persistent-exposure arm, at 0.296, still beats the no-playbook arm, at 0.251 (Table 4, reported). Exposure does not make the lessons worthless. It makes them worth less than the same lessons delivered on a trigger.
The same pattern shows up on the intervention side. Remove Sentry's intervention-need gate, so that guidance is generated at every scheduled monitoring point, and WebShop reward falls to 0.136, below the unassisted agent's 0.170 (reported). Advice the agent did not need is not free.
The mechanism

Sentry is a layer around the agent loop, not a new agent. Everything in it is an LLM call to the same model as the agent (Qwen3.5-9B by default), at temperature 0 with structured outputs, except the schema check.
1. Detection: an LLM judge over a five-step window
At every step Sentry looks at the task objective, the current action and the last reasoning-action-observation cycles. It first checks the action against the environment's schema. That check is mechanical: parsed, schema-valid, accepted by the environment. The README's integration sketch passes exactly those three flags (parsed_ok, schema_valid, accepted_by_environment) per step.
If the action is valid, an LLM detector prompt classifies the window against a fixed taxonomy and returns whether intervention is needed, a failure type, retrieval labels, local evidence and a short diagnosis. There are no hand-written heuristics per failure type in the paper. The taxonomy is:
| Failure type | Retrieval labels | Route |
|---|---|---|
| Action validity | malformed call, missing arguments, bad types, unsupported tool, invalid target | hard repair |
| Progress | repetition or looping · planning stalls · over-exploration · termination miscalibration | soft repair |
| Reasoning & grounding | hallucination · objective drift · reasoning-action mismatch | soft repair |
If nothing is wrong, nothing happens: no retrieval, no message, no lessons. That is the whole point.
2. Hard repair: the cheap case
An invalid action goes back to the agent with the action, the schema error, the schema and the recent trajectory, and the instruction to output only a corrected action. Verification is immediate: the corrected action either passes the schema and is accepted by the environment, or it does not. Hard repairs never write playbook entries.
3. Soft repair: retrieval keyed by the label
For progress and grounding failures, Sentry builds a soft-repair prompt from the failure type, the labels, the local evidence, general advice for that type, and up to playbook entries. Retrieval is deliberately not embedding search. It is a deterministic filter:
- keep entries with the same failure type and at least one shared label;
- rank by the number of shared labels, breaking ties by recency;
- take the top five, with at most two per primary label where possible, so one subtype cannot fill the slot;
- if nothing shares a label, fall back to the five most recent entries of the same type.
Because retrieval is capped, the playbook can grow without the prompt growing.
4. Verification, then the write
This is the rule that separates Sentry from a runtime monitor with a memory bolted on. After a soft repair, Sentry waits cycles and asks a verifier prompt one question: did the agent stop the diagnosed pattern and resume valid, task-relevant, observation-grounded progress? The verifier sees only the triggering failure and the ten cycles after it. It never sees the benchmark reward or whether the task eventually succeeded.
Only a repair judged resolved becomes a playbook entry, summarised as failure type, labels, trigger pattern and repair principle. Unresolved repairs write nothing, because a lesson learned from a failed repair would misfire later too. Detection and retrieval stay fixed; only the playbook evolves.
Here is what comes out, verbatim from the paper's example WebShop playbook:
[explore-00008] Over-exploration.
When an item page already shows enough evidence for the requested
constraints, commit instead of opening more unrelated candidates.
[term-00011] Termination miscalibration.
If Buy Now is visible and the observed item satisfies the hard
constraints, use click[Buy Now] instead of searching or navigating away.
[hall-00016] Hallucination.
Do not claim "no matching item exists" after only checking one
irrelevant result page.Compare hall-00016 with ACE's "no match on page 1, check the next pages". They are nearly the same advice. The difference is that Sentry's version is only shown to an agent the detector has just caught claiming nothing matches.

Results
Against runtime interventions

The setup: Qwen3.5-9B as the task agent and as every auxiliary model in every method, so differences come from the method and not the judge. All calls go through the Together API. Task sets are sampled once, then held fixed over five runs with shuffled order. WebShop, AppWorld and Mind2Web Replay use 100 tasks each, SWE-bench Lite 50. Sentry starts every run with an empty playbook and learns online, so its early tasks get only the general repair advice.
Metrics differ by benchmark: average reward on WebShop; test-pass rate (percentage of checks passed per task) on AppWorld; issue-resolution accuracy under the official harness on SWE-bench Lite; and on Mind2Web Replay, the mean fraction of reference steps completed in order.
What "+37% on average" averages. It is a relative gain, computed per benchmark against whichever baseline is strongest on that benchmark, then averaged without weights: (77.3 + 39.8 + 25.0 + 7.1) / 4 = 37.3% (reasoned, and it matches). The reference changes per column. It is Wink on WebShop and AppWorld, AgentForesight on SWE-bench Lite and AgentFixer on Mind2Web Replay. That is the conservative way to state it. Against any single fixed baseline the average is larger: 58.7% against AgentGuard, 70.7% against Wink (reasoned). The widget below runs the arithmetic.
Three things are worth reading off the table rather than the headline.
The WebShop gain is mostly about buying. The paper's full WebShop table counts tasks where the agent made any purchase: 23 for the base agent, 79 for Sentry. Full successes go from 6 to 14 (reported). WebShop gives partial reward for a purchase that matches some attributes, so a 9B agent that never commits leaves most of the reward on the table. Termination-miscalibration and over-exploration lessons are aimed at exactly that, though the paper does not break the gain down by label. The +77.3% is real, but it says more about how often Qwen3.5-9B fails to finish a shopping task than about deep reasoning repair.
SWE-bench Lite is small. It has 50 tasks, so accuracy 0.30 against 0.24 is about 15 resolved issues against 12 (reasoned, from means over five runs). Sentry intervened 90 times there, on only 9 guided tasks, and 5 of those were recovered (reported). The +25% is three issues.
Mind2Web Replay is the authors' own adaptation. In the original Mind2Web, each step is scored independently given the reference history. In Replay, a rejected action leaves the agent at the same recorded state with feedback and lets it retry, so the benchmark measures recovery by construction. The paper says plainly that its scores are not comparable to official Mind2Web numbers. It is also where Sentry's margin is smallest, +7.1% over AgentFixer.
With GPT-OSS-120B as both agent and judge, on WebShop and Mind2Web Replay only, the paper reports +51% on average over the best runtime baseline (95% and 8%).
Against context evolution, and combined with it

This table uses a different protocol, which the thread itself flags. Every method initialises its memory on 50 tasks and is scored on 50 disjoint tasks, still updating as it goes. The base agent scores 0.186 here against 0.170 in Table 1 because the task set is different. Do not subtract across the two tables.
"+39% over ACE" is again a mean of relative gains, over two benchmarks this time: (47.8 + 29.7) / 2 = 38.75% (reasoned). ACE + Sentry runs both memories separately: ACE's task-level playbook goes into the initial context as usual, and Sentry's recovery playbook stays outside and is consulted only on a diagnosed failure. The combination scores 0.555 and 0.789, the best in the table. That is +12% and +8% over Sentry alone (reasoned), and +66% and +40% over ACE alone, as the thread's chart puts it.
That composition is the paper's most practical result: task-level advice belongs in the prompt, and failure-level advice belongs behind a trigger.
Does the playbook carry over?
On held-out tasks, a frozen playbook learned on the 50 initialisation tasks lifts WebShop from 0.251 to 0.364 and Mind2Web Replay from 0.533 to 0.651, compared with the same Sentry machinery with retrieval disabled. Letting it keep learning during evaluation reaches 0.495 and 0.729 (all reported). The lessons transfer, and continued verified updates add more.
What each piece is worth
The component ablations (WebShop and AppWorld, online protocol) remove one piece at a time:
| Removed | WebShop | AppWorld | Mean relative drop |
|---|---|---|---|
| nothing (full Sentry) | 0.468 | 38.53 | — |
| hard repair | 0.297 | 31.20 | 27.8% |
| soft repair | 0.272 | 26.85 | 36.1% |
| intervention-need check | 0.136 | 23.70 | 54.7% |
| playbook self-evolution | 0.295 | 28.90 | 31% |
| failure-specific retrieval | 0.316 | 31.05 | 26% |
| recovery verification | 0.343 | 33.50 | 20% |
Scores are reported. The mean drops for the last four rows are the paper's, and I recomputed all six from its per-benchmark percentages (reasoned). The gate is the biggest single component, which is the paper's thesis in one row: the most damaging thing you can do is intervene when nothing is wrong. Without hard repair, Sentry still beats every runtime-intervention baseline on both benchmarks (0.297 against Wink's 0.264, and 31.20 against 27.56), so the gains are not only format fixing.
Retrieval specificity matters in the order you would expect. On held-out tasks, relative to label-filtered : failure-type-only retrieval loses about 11.3%, unfiltered retrieval about 18.8%, random entries about 28.9% (reported, averaged over WebShop and Mind2Web Replay). The value of matters less: loses about 4.5% and about 6.2%. More lessons is not better even when they are the right kind.
How often repairs work, by whose judgement

Sentry resolves 81.7% of 939 detected failures (reported). Action-validity failures are easiest because acceptance is binary; progress failures on SWE-bench Lite are hardest, at 34.0%.
Two cautions. First, this rate is Sentry's verifier grading Sentry's repairs, with the same 9B model. Second, the text quotes category rates of 92.6%, 64.0% and 70.2%, while the figure's dashed lines read 87.6%, 60.6% and 64.4%. The dashed lines are the unweighted means of the four benchmark bars (reasoned: (97.8 + 90 + 69 + 93.7) / 4 = 87.6). The text's numbers are presumably pooled over all failures, weighted by count. The per-benchmark counts are not given, so I cannot reconcile the two exactly.
The paper does two things to check its own judge. A paired fork at 200 soft-repair checkpoints gives one branch Sentry's repair and lets the other continue unaided. Local recovery goes from 33.5% to 60.5%. On final score the repaired branch wins 74 pairs, loses 25, and ties 101 (reported). Those 25 losses are worth remembering: a correct-looking hint sometimes leaves the agent worse off. A human audit of 200 trajectory windows (113 interventions plus 87 matched controls) finds 86.0% agreement on whether to intervene, 81.8% on whether the repair recovered, and 85.8% of interventions judged locally helpful, with Cohen's between 0.82 and 0.89 across annotators. The annotators are members of the research team, which the paper states.
What it costs
Across the four benchmarks, Sentry uses 19.01M tokens against the base agent's 12.32M, which is 1.54x. AgentForesight uses 2.44x and Wink only 1.01x (reported). Wall-clock time moves both ways. Sentry adds 7.1 s per task on WebShop and 46.0 s on Mind2Web Replay, but saves 31.7 s on SWE-bench Lite and 45.6 s on AppWorld, where breaking a loop early saves more than the extra judge calls cost.
What I would take from it
- The finding is bigger than the method. "Conditional knowledge should be conditionally exposed" applies to any agent memory: Hindsight's reflect step, TencentDB's tiered memory, the playbooks in Lilian Weng's harness taxonomy. MemHarness reached a neighbouring conclusion from the other side: a retrieved memory replayed verbatim can hurt more than no memory, so it rewrites memories against the current state first. Sentry changes when a lesson appears. MemHarness changes what the lesson says when it does.
- The write rule is the part to steal. "Only store a lesson if the fix worked, judged without the reward" is cheap to add to any reflection loop, and the ablation puts it at about a 20% mean drop when removed. It is the same instinct as the noise-adjusted acceptance floor in RRSI: do not let a self-improving loop learn from its own noise. Dream-RSI, whose outer loop reuses a cache of past attempts, faces the same admission question.
- The detector is the risk. Everything hinges on a 9B judge reading a five-step window; when it misfires, Sentry is AgentGuard with better prose. Removing the gate drops WebShop below the unassisted agent, so its precision is load-bearing. The paper itself lists the shared model for detector, verifier and agent as a limitation.
- The playbook will need garbage collection. The paper says entries can become redundant or overly specific and proposes merging and pruning. The bounded keeps the prompt flat, not the advice fresh.
What I could and could not check
Every benchmark figure above is the paper's; I re-derived the headline averages, the ablation means, the token ratio and the Figure 3 dashed lines from its own tables, and they agree. What I could not check is the method itself. At repo commit 83064d1 (measured) there is no code: the README's quick start (sentry-validate configs/paper.yaml, sentry run appworld) names files that are not in the repository yet. The demo video is labelled as a recorded GPT-4.1 run on AppWorld, a model the paper never evaluates, so I read it as an illustration, not evidence.
Sources: the paper (arXiv 2610.02994), the repository, the authors' thread on X, and for background ACE and AgentGuard.