FAR lets world models learn which memories to recall from future supervision, beating fixed retrieval rules across LoopNav, SoundSpaces, and AI2-THOR
Synopsis
The work proposes Future-Aware Recall (FAR): during training it measures predictive utility by the negative diffusion prediction loss of the realized future under a video diffusion world model, uses that signal to supervise a retriever that stays future-blind at inference, and learns both cue-specific relevance and query-dependent weights over cues such as time, pose, vision, and audio, outperforming hand-designed recall rules across LoopNav, SoundSpaces, and AI2-THOR.
Interpretation
FAR turns 'which past to recall' into a learnable problem supervised by the future: at training time predictive utility is approximated by the negative diffusion prediction loss of the realized future conditioned on a candidate memory, and a Bayesian posterior transfers that credit to a retriever that cannot see the future at inference. Prior external episodic memory for world models relied on fixed relevance criteria such as temporal recency, pose or field-of-view overlap, or visual embedding similarity; FAR instead optimizes the retrieval rule itself against the downstream prediction objective, supported by a predictive relevance bound and a latent-variable maximum-likelihood derivation. The paper provides the derivation of Proposition 1, a latent-context maximum-likelihood view, and a candidate-level EMDR2-style approximation; in the video diffusion instantiation predictive utility is approximated by negative prediction loss averaged over four diffusion timesteps with shared perturbations across candidates.
FAR learns a cue-specific relevance score for each available cue and fuses them with query-dependent gating weights, so which cue to trust can change from query to query. Existing external retrievers typically use a single cue or a fixed combination; FAR's fusion weights are produced from the current query's cue statistics, and cues are used only to select memories rather than being passed to the world model as generation conditions. On SoundSpaces, metadata alone is a strong signal with periodic scans, whereas under sparser endpoint scans the multi-cue advantage grows with return-phase length, particularly in LPIPS and DreamSim; in the elbow-corridor example, audio retrieves a clearer view when pose and visual geometry are ambiguous.
In environments whose state changes, FAR recalls history consistent with the current world state and reduces errors caused by stale memories. Retrieval driven mainly by recency or geometry can pick a spatially well-matched but outdated memory; FAR ranks memories by how well they support accurate prediction. In AI2-THOR, accuracy is stratified by state-change magnitude, and multi-cue FAR outperforms temporal and geometry-based recall for both surface changes and container reveals, with the advantage especially pronounced for container reveals; in a paired counterfactual example, the same current observation and action yield different futures depending on whether a tomato was previously placed in the refrigerator, while Temporal and WorldMem fail to preserve that dependence.
FAR also supports off-scene dynamics prediction: in a two-agent corridor setting, FAR with only time and pose metadata exceeds the temporal and WorldMem baselines, and adding an observation cue improves it further. This indicates future-aware training can learn when informative crossings are likely to occur, not merely which location matches geometrically. The paper reports accuracy values for Temporal and WorldMem, a higher value for FAR with metadata alone, and a further gain for the multi-cue variant; each recalled context consists of two observations separated by several seconds to capture motion, and evaluation uses a one-step both-ends probe rather than an autoregressive rollout.
Perspective
The work focuses on reading from an external episodic memory, that is, choosing which past observations enter a compact context; it targets world-model research that needs long-horizon persistence and is validated in controlled simulations (LoopNav in Minecraft, SoundSpaces on Matterport3D, AI2-THOR in iTHOR), with the available cue set varying by environment. The method is agnostic to cue type, so temporal, spatial, visual, and auditory signals can all be used, and retrieval cues only rank memories rather than conditioning generation, which makes it easy to combine with existing generator interfaces. The authors describe it as complementary to internal long-context memory and persistent-state representations.
The authors list several open directions: memory writing, compression, forgetting, and higher-order interactions among recalled memories are left to future work; the method requires informative retrieval cues and adds training-time computation; the diffusion-loss utility may underweight semantically important local changes, and alternative task-aware utility functions are a promising direction; experiments use controlled simulations, motivating evaluation in richer real-world settings and integration with learned memory formation and persistent-state representations. In addition, the candidate-level approximation affects only retrieval credit assignment and does not explicitly score memories that become informative only when recalled together, though such interactions remain represented by the set-level model. Readers of the abstract and main text alone should consult the figures and tables for specific numeric values.
