Skip to main content
Back to timeline
arXivSource publication:

Alibaba team's PoS adds explicit belief states plus Belief Trapping detection and recovery, topping all four benchmarks across three backbones

Synopsis

The work introduces PoS, an inference-time framework that replaces interaction history as the agent's decision context with a continually maintained explicit belief state (world state plus unresolved epistemic and achievement gaps), validates belief updates through a Belief Sentinel, detects Belief Trapping (stagnation, cycles, drift) from window-level signals, and composes recovery constraints tailored to the trapping pattern and blocked gap type; across four benchmarks (ALFWorld, LOCA-Bench, RCA-100, ClinDiag) and three backbones (Qwen3.7-Plus, Kimi-K3, GLM-5.3), PoS achieves the highest overall performance in all 12 benchmark-backbone combinations, with relative gains over the strongest same-backbone baseline reaching 22.68% on ALFWorld and 37.89% on RCA-100 joint accuracy.

AI-generated editorial illustration: Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States

Interpretation

PoS replaces the agent's decision context with a task-conditioned explicit belief: the world state is represented as an Entity-State-Relation structure annotated with provenance and confidence, while unresolved requirements are split into epistemic gaps (what still needs to be learned) and achievement gaps (what still needs to be accomplished), with one active gap selected as the current focus. Memory-based methods (Raw Trajectory, ACON, PACE, HiAgent) produce transformed records of historical evidence, and state-oriented methods such as LongHorizon-Harness rely on independently verifiable intermediate artifacts; PoS instead merges 'what the world currently is' and 'what remains unresolved' into one inspectable, revisable decision state. The paper provides a POMDP formalization, belief-update equations, and an active-gap selection mechanism, with a full notation table and algorithm in the appendix; this is a constructive methodological contribution whose value is supported indirectly by the four-benchmark experiments.

PoS introduces a Belief Sentinel that validates candidate belief updates for consistency, distinguishing internal inconsistency (mutually conflicting states assigned to the same entity) from external inconsistency (contradiction with the latest observation or other interaction evidence), after which the task agent revises against the evidence before committing. The paper notes that belief updates are generated by the LLM itself and may inherit erroneous assumptions or hallucinations; earlier work built explicit beliefs but did not make continual validation a fixed part of belief maintenance. Ablations show that removing consistency validation lowers performance on all four benchmarks, with the largest losses on ALFWorld and LOCA-Bench and smaller losses on ClinDiag; the paper attributes this pattern to the greater risk of inconsistent updates when complex interactions require reconciling changing states and observations.

The paper names and characterizes Belief Trapping, where the agent keeps acting without meaningfully advancing toward the goal, and estimates belief health from three window-level signals (gap persistence, progress stagnation, belief recurrence); after detection it performs factorized diagnosis along agent dynamics (static, cycle, drift) and blocked gap type, then composes pattern-specific escape constraints with gap-specific progress constraints. T3 only detects belief deviation and truncates uninformative trajectories during reinforcement learning, and CGDP uses an exhaustion gate to terminate unproductive context gathering; neither restores progress during inference, whereas PoS keeps the active gap unchanged and applies additional constraints to redirect subsequent actions. Ablations show that removing Trapping Diagnosis and its recovery degrades performance on all four benchmarks, with larger losses on RCA-100 and ClinDiag; against a variant that replaces factorized recovery with a generic recovery prompt, PoS's recovery is better on all four benchmarks.

Across all 12 benchmark-backbone combinations, PoS achieves the highest overall performance, with relative gains over the strongest same-backbone baseline reaching 22.68% on ALFWorld, on LOCA-Bench, 37.89% on RCA-100 joint accuracy, and on ClinDiag; on LOCA-Bench, as context grows from 8K to 256K, PoS remains broadly stable from 96K to 256K and exceeds the strongest baseline by several points at 256K. The paper also reports reversals among baselines: PACE falls below Raw Trajectory on LOCA-Bench, LongHorizon-Harness delivers less consistent gains on RCA-100 and ClinDiag, and on ClinDiag with GLM-5.3 none of the context-management baselines outperforms Raw Trajectory; PoS's gains span both execution and diagnosis tasks. The main table covers 134 ALFWorld cases, 525 LOCA-Bench cases, 103 RCA-100 cases, and 604 ClinDiag cases (302 Common and 302 Rare, with the Emergency subset excluded due to data-access constraints), with all methods using the same benchmark adapters, tool interfaces, and action budgets, and all LLM calls using deterministic decoding at temperature 0.

Perspective

The framework targets long-horizon agent decision-making that requires maintaining world state across many steps, covering execution tasks (ALFWorld, LOCA-Bench) and evidence-seeking diagnosis tasks (RCA-100, ClinDiag), and it runs at inference time without additional training, so it can be layered onto existing backbones. The evaluation spans three backbones and four benchmarks, with all methods sharing the same benchmark adapters, tool interfaces, termination protocol, and action budgets, and with deterministic decoding for LLM calls. For a reader, this means PoS is best understood as a replacement for the context-management layer: it does not change underlying model capability but changes the state representation each decision is conditioned on. The paper also notes that explicit belief maintenance adds token and computational cost, since natural-language beliefs require repeated text processing, and suggests future work on task-conditioned latent world-state representations to reduce overhead; PoS also does not systematically integrate external domain-knowledge retrieval, so identifying relevant evidence and interpreting observations depends largely on the backbone's existing knowledge.

This is a fast-parse version, and the equation numbering, the specific values in Figure 3 and Figure 4, and some relative-gain percentages are not fully rendered in the text, so precise readings of trapping incidence distributions, recovery-strategy margins, and the context-scaling curves still require the original figures. In addition, Belief Trapping detection depends on a set of hyperparameters (window size, maximum recurrence lag, recurrence distance threshold, health threshold, and others); the paper gives their values but does not expand on their sensitivity in the main text. Belief health uses a weighted form that multiplies the stronger stagnation-or-recurrence signal by gap persistence, and the appendix illustrates its tradeoffs against geometric mean and maximum formulations through three boundary configurations, but behavior under different task distributions remains worth watching. On cost, PoS reduces Task Agent token consumption on RCA-100 while total consumption rises because of belief construction and maintenance, and the paper lists reducing this overhead as future work. Finally, the paper itself notes that explicit belief cannot replace execution skills or domain knowledge: success on LOCA-Bench's Env.Ops category remains in a low range, and many incorrect ClinDiag diagnoses occur after the agent has received correct information but lacks the medical knowledge to interpret it.

Sources