Skip to main content
Back to timeline
arXivSource publication:

ReFract benchmark: the same query demands different actions by user role, and top LLMs solve at most 69% while over half of trajectories attempt role-violating actions

Related research and updates

Synopsis

The authors introduce ReFract, a benchmark of 150 expert-validated entries that uses Text World Models to simulate industrial maintenance and equipment fault troubleshooting environments and requires an agent to act differently on the same query depending on the user's role; state-of-the-art LLMs solve at most 69% of tasks, and more than 50% of their trajectories contain attempts at perspective-violating actions.

Source-provided article image: ReFract: Benchmarking Perspective Awareness in Language Model Agents with Text World Models
Figure 1 ·

Figure 1 : Different roles asking the same query pursue different goals . A perspective-aware LLM agent ought to infer the goal of the given role and take actions admissible inside the role’s capability and knowledge boundary. In cases where the role does not have the remit to accomplish the goal, the agent must identify capable roles and route the work accordingly.

arXiv

Interpretation

The paper proposes and names Perspective Awareness: the agent must infer what a role intends and act only through tools that role may legitimately use. Existing benchmarks largely overlook inferring role intent and constraining tool permissions; this work isolates that axis and makes it the object of evaluation. The capability definition and motivation come from the abstract's discussion of high-stakes settings such as industrial maintenance and equipment fault troubleshooting; it is a conceptual contribution without quantitative validation.

The authors build ReFract, a benchmark of 150 expert-validated entries whose core design is that the same query requires different actions under different user roles. Entries are grounded in anonymized queries from domain support conversations, against which Text World Models simulate the agent's operating environments and perspective-aware action trajectories are assembled. The abstract states the entry count (150) and the expert-validated provenance, plus the Text World Model and trajectory-assembly method; it does not report entry distribution or annotation-agreement details.

On state-of-the-art LLMs, at most 69% of tasks are solved, and more than 50% of trajectories contain attempts at perspective-violating actions. This presents perspective awareness as a distinct, largely unsolved axis of agent evaluation rather than a capability already covered by existing benchmarks. The abstract reports two quantitative indicators, the task-solve ceiling and the share of violating trajectories, but does not name the evaluated models, the number of evaluation runs, or statistical uncertainty.

Perspective

The work targets LLM agents that must act within role-permission boundaries, in high-stakes settings such as industrial maintenance and equipment fault troubleshooting where actions are enacted on physical equipment; because the evaluation vehicle is Text World Models, the conclusions apply to assessing role calibration in text-simulated environments. For researchers and deployers building role-aware agents, ReFract offers a reusable task-construction approach (anonymized domain support queries plus Text World Models plus perspective-aware trajectories) and quantitative baselines (at most 69% solve rate, over 50% violating trajectories).

The abstract does not list the evaluated models, the number of evaluation runs, or statistical uncertainty, nor does it describe how the 150 entries distribute across roles and scenarios, so the robustness of the 69% and 50% figures still needs confirmation from the full text. The fidelity of the Text World Models to real equipment and permission systems, and the specific expert-validation process and its consistency, are open questions a reader would want to check. The abstract also does not distinguish attempted perspective-violating actions from actual harm.

Sources