EvolvingNav predicts where targets go with a time-indexed 4D belief, lifting first-inspection success from 45.33% to 61.32% on EvoWorld-Bench
Synopsis
The work introduces EvolvingNav, which builds a persistence–relocation belief from timestamped 3D object histories and closes the loop with an event-driven predict–observe–replan filter so an agent can infer where a target is when targets may move before the query or during navigation, and it releases EvoWorld-Bench with 54 scenes and 803,680 tasks, reporting improved navigation success and search efficiency over the evaluated baselines in simulation and real-robot experiments.
Interpretation
The paper formalizes Evolving-World Navigation: the agent has only a causal observation history up to query time, the target may move before the query and may keep moving during execution, so the agent must estimate whether the target will still be at a candidate when it arrives. Prior work studied temporal prediction and navigation with moving objects separately, for example PredictiveGraphs predicting future object–receptacle states and updating estimates during navigation, and Transit-Aware Planning considering portable targets that move while the agent travels; this work combines irregular observation histories, candidate-specific arrival times, and visibility-conditioned evidence within one navigation loop. The problem setting is formalized in Section 1 and Section 3.1, and distinguishes fixed-target episodes from continuing-evolution episodes.
EvolvingNav builds a time-indexed belief through a persistence–relocation decomposition: it predicts the probability that the target stays at its last observed state, scores every alternative with a shared pointer head, and retains probability mass outside the known candidate set. Unlike a Direct Transformer that normalizes candidate logits directly, this structured factorization separates staying from relocating; ablations show that removing predictive belief drops First-Inspection SR from 62.17% to 12.51%, and direct candidate prediction reduces First-Inspection/Search SR by 13.84/15.75 points. On prediction quality, the predictor improves Top-1 by 4.08 points on the temporal split and 9.47 points on leave-one-scene-out relative to the Direct Transformer, with NLL decreasing by 0.2367 and 0.3379 respectively.
An event-driven predict–observe–replan filter forecasts occupancy at each candidate's estimated arrival time and applies visibility-conditioned soft updates: a missed detection downweights a location hypothesis by a calibrated likelihood rather than deleting it, and time-valid evidence rounds prevent reusing the same observations. Ablations show Hard Removal degrades rank, Top-1, NLL, and ECE, while the calibrated Bayesian update achieves the best rank, Top-1, and ECE; removing evidence updates reduces Search SR by 42.02 points and raises inspections to 6.48. The text gives a concrete update example: causal memory initially assigns probability 0.62 to the last-seen coffee table, the calibrator yields a detection probability such that the likelihood at that state is 0.069 when the target is not detected, the posterior shifts accordingly, and the inspected state retains nonzero mass.
The paper releases EvoWorld-Bench, grounded in human activity traces, with 54 scenes and 803,680 task instances, and uses paired static, routine, and random worlds that hold scenes and queries fixed while changing only the hidden transition mechanism. The paired control separates using historical regularity from using category frequency or spatial convenience: the SR gain over Last Seen is 22.36 points in routine worlds but only 6.93 points under random transitions, locating the advantage in learnable temporal regularity. Benchmark construction distinguishes measured, environment-derived, simulated, and benchmark-authored quantities, and audits temporal and metadata leakage, split overlap, paired-world consistency, physical validity, and patrol access to hidden state.
Perspective
The result targets embodied agents that must find movable targets across repeated visits, in settings such as households where human activity has regularities and targets may move before the query or during navigation; the method relies on timestamped 3D observation histories, a legal candidate set, reachable inspection viewpoints, and a frozen zero-shot vision-language controller that invokes memory, prediction, inspection, exploration, and navigation tools. For readers, this means the arrival-time state prediction can be treated as an output interface of a memory system rather than storing only the last observed location; the real-robot part uses the DEEP Robotics LYNX M20 as the primary platform, with X30 and Lite3 on transfer subsets, indicating the belief interface can be reused across platforms.
The text states that future work will address interacting objects and continuously changing goals, so the current setting does not yet cover those cases. Paired experiments show the advantage concentrates on learnable temporal regularity: under random transitions the gain over Last Seen falls to 6.93 points and the method trails Transformer Direct by 1.27 points; in cross-environment results Uni-LaViRA leads on R2R-CE SR and NaVILA leads on SAGE-Bench SPL, so the conclusion should not be generalized to uniform superiority across all navigation protocols. The behavior-shift diagnostic shows Online Recovery SR falls for all methods under phase offsets, return delays, destination shifts, and their combination, and the authors explicitly do not present it as evidence of unrestricted out-of-distribution generalization. In addition, real-robot success requires autonomous online confirmation within three inspections during a 360-second episode, with missed detections, navigation failures, timeouts, and human interventions counted as failures, which affects how results are read.
