HIDE benchmark and SEEK memory framework: 62.9% average success across 15 hidden-state tasks, 11.7 points above the strongest baseline
Synopsis
The authors introduce HIDE, an RLBench-based benchmark of 15 manipulation-memory tasks spanning repetition counting, historical-state recall, and execution-progress tracking under partial observability, together with SEEK, a framework combining windowed context memory, persistent anchor memory, and stage-counter memory to retain historical evidence and track execution state; evaluations show substantial limitations in existing policies on HIDE, while SEEK reaches 62.9% average success in simulation, 11.7 percentage points above the strongest baseline SAM2Act+, and 89% average success on four real-robot tasks.
Interpretation
The paper formulates manipulation memory around hidden task states: there exist two interaction histories that lead to visually equivalent current observations yet require different optimal actions, and builds HIDE on this definition with 15 tasks in three categories—repetition counting, historical-state recall, and execution-progress tracking—five tasks each. Prior benchmarks such as RLBench, CALVIN, LIBERO, and RoboCasa focus on task completion and generalization, while memory-oriented benchmarks including MemoryBench, MIKASA-Robo, RoboMME, RoboMemArena, RMBench, and LIBERO-Mem measure memory demand through task length, scene complexity, or the amount of retained context; HIDE instead organizes tasks by the hidden variables required for correct decisions and explicitly constructs decision points where the current observation is insufficient and history supplies the evidence. Built on RLBench with randomized configurations and appearance variations, automated demonstration generation, and structured hidden-state annotations; each task uses 100 training demonstrations and 25 held-out test episodes, and difficulty can be varied through the number of repetitions, the number and similarity of distractors, the delay between informative observations and later decisions, and the number of manipulation stages.
SEEK combines three complementary memory mechanisms—Windowed Context Memory (WCM) for recent interactions, Persistent Anchor Memory (PAM) that retrieves a historical entry after it leaves the window based on the current observation, and Stage-Counter Memory (SCM) representing progress with a discrete counter and a learned stage embedding—read through SAM 2-style attention with a bounded budget of at most six recent entries, one anchor, and one stage representation. Existing memory policies retain history through temporal windows, recurrent states, or memory banks, but the paper argues that retaining history does not ensure recovery of the required hidden state; SEEK splits memory responsibilities according to explicit hidden-state requirements and separates WCM/PAM, which describe previous interactions, from SCM, which indicates how far execution has progressed. SEEK builds on a language-conditioned coarse-to-fine multi-view policy: the coarse branch uses memory-conditioned features to predict workspace translation heatmaps, and the fine branch refines the selected interaction using current local geometry; memory is written after action prediction, only preceding observations are eligible for retrieval, and memories and the counter reset between episodes; training is two-stage behavior cloning, with Stage 1 learning a memory-free policy and adapting the visual backbone via LoRA, and Stage 2 freezing the visual backbone while training memory and stage-related parameters.
On HIDE, existing policies remain limited and their strengths diverge by category: SAM2Act+ raises average success from 42.7% for SAM2Act to 51.2% but reaches only 46.4% on repetition counting and historical-state recall, while MME reaches 35.2% on historical-state recall but only 2.4% on repetition counting; SEEK attains 62.9% average success with 61.6%, 59.2%, and 68.0% across the three categories, corresponding to gains of 15.2, 9.6, and 7.2 points over the strongest category-wise baselines. The paper compares vision-language-action models, 3D manipulation policies, and memory-based methods on one benchmark, showing that the strongest baseline varies by category—SAM2Act+ leads repetition counting and execution-progress tracking, while GR00T-N1.7 leads historical-state recall at 49.6%—whereas SEEK leads all three. A single policy is jointly trained across all 15 tasks, with per-task success rates, category averages, and an overall average over 15 tasks; ablations show that storing visual features rather than proprioception raises average success from 42.5% to 62.9% (repetition counting +30.4, historical-state recall +17.6, execution-progress tracking +13.4), and that increasing memory length from 0 to 6 raises average success from 46.7% to 62.9%, with repetition counting dipping to 44.2% at length 2 before rising to 61.6% with longer context.
Beyond HIDE, SEEK retains strong performance: 84.7% average success on the standard 18-task RLBench, slightly above its SAM2Act reproduction at 84.1% and above RVT-2 at 81.4% and RVT at 62.9%; on The Colosseum it reaches 68.6% on clean scenes and 61.9% averaged over perturbations, a 9.8% relative drop; and on four real-robot tasks it reaches 89% average success versus 47% for SAM2Act+. The paper tests the memory design on standard manipulation, environmental perturbations, and physical execution, indicating that adding memory preserves standard-task performance rather than trading it away, and provides real-robot comparisons. Real-robot evaluation uses a Franka Emika Panda with a Robotiq gripper and an Intel RealSense D455 camera, collecting 50 demonstrations per task and evaluating over 25 test trials with task variations uniformly represented in collection and evaluation; the four tasks are pressing a button a specified number of times, stacking a specified number of cups, cleaning a desk, and lifting blocks to search for a white piece.
Perspective
This work targets manipulation settings where hidden states must be recovered from interaction history: repetition counting, historical-state recall, and execution-progress tracking, evaluated in RLBench-derived simulation and in a controlled indoor real-robot setting with button pressing, cup stacking, desk cleaning, and searching for a hidden white piece. For researchers who want to evaluate or design memory mechanisms, HIDE offers a controlled platform with structured hidden-state annotations, randomized configurations, and automated demonstration generation, while SEEK offers a composable template of recent context, persistent anchor, and execution progress, showing that the memory read budget can be bounded regardless of archive length. The paper explicitly scopes itself to controlled and interpretable forms of hidden state and names richer real-world settings, broader sources of partial observability, and more general memory mechanisms as future directions.
Open questions remain: HIDE currently focuses on controlled and interpretable hidden states, so whether richer real-world sources of partial observability show the same division of labor is untested; whether the three mechanisms and the read budget still apply outside the paper's task distribution needs more validation; the real-robot results rest on four tasks with 25 test trials each, and although task variations are uniformly represented in collection and evaluation, the task variety is limited; and in the text available here the per-perturbation breakdown tables for The Colosseum are empty, so robustness can only be read from the clean-scene and perturbation-average summary numbers.
