Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

LiteTrajEval pairs budget-bounded rule profiles with a single rubric-guided LLM judge, improving failure-localization alignment with human annotations by roughly 20–35 percentage points while cutting cost about 6x and evaluation time more than 8x

The work presents LiteTrajEval, a lightweight architecture for budget-bounded trajectory evaluation that derives compact domain-specific rule profiles offline, then preprocesses each trajectory online, marks heuristic failure signals, serializes it under a fixed global budget, and invokes a single rubric-guided LLM judge to produce structured diagnostic reports; on public Magentic-One-style and tau-bench-style trajectory datasets it improves failure-localization alignment with human annotations by roughly 20–35 percentage points on Magentic-One and up to 23 percentage points on tau-retail compared with AgentRx, while reducing cost by about 6x and evaluation time by more than 8x, and it has also been deployed in the authors' enterprise agentic platform.