Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

ReFract benchmark: the same query demands different actions by user role, and top LLMs solve at most 69% while over half of trajectories attempt role-violating actions

The authors introduce ReFract, a benchmark of 150 expert-validated entries that uses Text World Models to simulate industrial maintenance and equipment fault troubleshooting environments and requires an agent to act differently on the same query depending on the user's role; state-of-the-art LLMs solve at most 69% of tasks, and more than 50% of their trajectories contain attempts at perspective-violating actions.