Skip to main content
Back to timeline
arXivSource publication:

Remory adds residual memory tokens to summaries, lifting SummHay citation F1 by 4.04 points and approaching full-context joint score

Synopsis

Remory trains a neural memory network to append a bounded sequence of soft memory tokens after a generated summary, helping a frozen LLM approximate the continuation it would produce with the full history; on SummHay it raises citation F1 by 4.04 points and joint score by 3.55 using about 5.2% of input positions, and yields consistent gains for Qwen3.8-27B and GLM-5.3-Flash on long-horizon agent benchmarks while reducing repeated tool outputs and tool errors.

Source-provided article image: REMORY: Learning Residual Memory for Context Compaction
Figure 1 ·

Figure 1: BrowseComp and Terminal-Bench 2.1 scores. Pale and solid bars compare the same local actor without and with residual memory. Gray bars are published frontier references, with their original harnesses and protocols ( Google DeepMind, 2026 ; Terminal-Bench Team, 2026 ; OpenAI, 2026b ) ; they are not controlled comparisons with our runs.

arXiv

Interpretation

Remory treats the compacted summary as a readable checkpoint and appends a bounded sequence of continuous soft memory tokens after it, forming an analogue of a residual connection along the sequence dimension that helps a frozen actor approximate the continuation it would produce with the full history. Prior agent memory often combines summaries with retrieval, while continuous compression methods learn representations of longer text; Remory keeps the textual checkpoint and retrieval interface and learns an additional continuous representation conditioned on the checkpoint. The construction of summary and soft tokens, 64 soft tokens per block of up to 1,024 source positions, a final memory bounded at 4,096 slots, and memory networks with 1.94B and 1.502B trainable parameters for Qwen3.8-27B and GLM-5.3-Flash are stated in the text.

On SummHay, residual memory improves source attribution at nearly unchanged insight coverage: citation F1 rises by 4.04 points and joint score by 3.55, with a coverage change of 0.40. Compared with a text supplement given the same additional-position budget (summary++), the learned supplement preserves attribution cues more effectively; summary++ remains close to summary alone. Evaluation covers all 92 queries from ten collections; coverage and joint score average over all 621 reference insights, citation F1 over covered insights; on the 553 insights covered by both summary conditions citation F1 still rises from 57.31 to 61.41; a paired collection bootstrap gives a 95% interval for the joint-score gain, conditional on the generated and judged outputs.

On long-horizon agent benchmarks, residual memory yields consistent gains: Qwen improves by 9.8 points on AutomationBench, 7.6 on JobBench, 4.5 on Terminal-Bench, and 3.0 on BrowseComp; GLM gains 3.3 and 4.1 on Terminal-Bench and BrowseComp. The gains also hold for a stronger actor with a larger context window, indicating the interface is not tied to one model scale. Each within-actor comparison holds weights, the Codex CLI harness, prompts, tools, decoding, budgets, and compaction policy fixed, and both configurations can retrieve exact history; full task sets are used (1,266 BrowseComp and 89 Terminal-Bench tasks for both actors in one run; 600 AutomationBench and 65 JobBench tasks for Qwen in five runs), with task verifiers for Terminal-Bench and AutomationBench and matched GPT-5.6-Sol judges for BrowseComp and JobBench.

Residual memory reduces repeated tool outputs and tool errors and is preferred for the first action after compaction: across the four actor–benchmark pairs repeats fall by 17.9–29.8% and errors by 25.3–49.7%; among 488 paired next actions, residual actions receive 345 wins (70.7%), 21 ties (4.3%), and 122 losses (25.0%). These results link the memory channel to the first post-compaction decision and are consistent with better use of earlier feedback. Repeats count identical outputs from the same tool within a task, excluding planning and polling; errors count failed commands or explicit tool errors; the action-preference evaluation holds the persistent prompt and summary fixed, withholds candidate reasoning, and leaves counterfactual calls unexecuted, so it judges proposed actions rather than measuring outcomes.

Perspective

The work targets long-horizon agents that must keep working within a finite window: the summary provides a readable checkpoint, soft tokens carry additional predictive information, and retrieval supplies exact records when needed. The memory interface is trained on a frozen actor, so it applies where an actor and compaction policy already exist and where up to 4,096 residual slots can be appended at compaction. The SummHay evaluation runs without tools, archive access, or source retrieval to isolate the representation itself, while the end-to-end evaluations retain exact-history retrieval. Cost estimates price actor and summary tokens at September 2026 uncached provider rates, counting residual positions as input, and exclude source-feature encoding, memory-network execution, and tools, so they measure token-priced generation rather than total execution cost.

The 95% interval for the joint-score gain is conditional on the generated and judged outputs, so its range depends on the generation and judging process. The first-action preference evaluation judges proposed actions and does not measure the outcomes those counterfactual actions would have produced. In the Live Trading case the two trajectories differ in market selection and timing, limiting return comparisons, and compactions per 1,000 responses rise from 10.5 to 14.9 even as the total falls with less interaction. The cell-segmentation case illustrates complementary roles for summary, memory, and retrieval but does not establish which omitted details came from memory. Cost estimates exclude source-feature encoding, memory-network execution, and tools. In addition, several equations and figures (Equations 1–3, Figures 1–5) are not fully rendered in the parsed text, so readers who need the exact forms and curves should consult the original.

Sources