Public articles linked to the same research event.
arXiv The authors introduce ReFract, a benchmark of 150 expert-validated entries that uses Text World Models to simulate industrial maintenance and equipment fault troubleshooting environments and requires an agent to act differently on the same query depending on the user's role; state-of-the-art LLMs solve at most 69% of tasks, and more than 50% of their trajectories contain attempts at perspective-violating actions.
The authors introduce ReFract, a benchmark of 150 expert-validated entries that uses Text World Models to simulate industrial maintenance and equipment fault troubleshooting environments and requires an agent to act differently on the same query depending on the user's role; state-of-the-art LLMs solve at most 69% of tasks, and more than 50% of their trajectories contain attempts at perspective-violating actions.
The authors introduce ReFract, a benchmark of 150 expert-validated entries that uses Text World Models to simulate industrial maintenance and equipment fault troubleshooting environments and requires an agent to act differently on the same query depending on the user's role; state-of-the-art LLMs solve at most 69% of tasks, and more than 50% of their trajectories contain attempts at perspective-violating actions.
The authors introduce ReFract, a benchmark of 150 expert-validated entries that uses Text World Models to simulate industrial maintenance and equipment fault troubleshooting environments and requires an agent to act differently on the same query depending on the user's role; state-of-the-art LLMs solve at most 69% of tasks, and more than 50% of their trajectories contain attempts at perspective-violating actions.