Public articles linked to the same research event.
arXiv The work presents LiteTrajEval, a lightweight architecture for budget-bounded trajectory evaluation that derives compact domain-specific rule profiles offline, then preprocesses each trajectory online, marks heuristic failure signals, serializes it under a fixed global budget, and invokes a single rubric-guided LLM judge to produce structured diagnostic reports; on public Magentic-One-style and tau-bench-style trajectory datasets it improves failure-localization alignment with human annotations by roughly 20–35 percentage points on Magentic-One and up to 23 percentage points on tau-retail compared with AgentRx, while reducing cost by about 6x and evaluation time by more than 8x, and it has also been deployed in the authors' enterprise agentic platform.
The work presents LiteTrajEval, a lightweight architecture for budget-bounded trajectory evaluation that derives compact domain-specific rule profiles offline, then preprocesses each trajectory online, marks heuristic failure signals, serializes it under a fixed global budget, and invokes a single rubric-guided LLM judge to produce structured diagnostic reports; on public Magentic-One-style and tau-bench-style trajectory datasets it improves failure-localization alignment with human annotations by roughly 20–35 percentage points on Magentic-One and up to 23 percentage points on tau-retail compared with AgentRx, while reducing cost by about 6x and evaluation time by more than 8x, and it has also been deployed in the authors' enterprise agentic platform.
The work presents LiteTrajEval, a lightweight architecture for budget-bounded trajectory evaluation that derives compact domain-specific rule profiles offline, then preprocesses each trajectory online, marks heuristic failure signals, serializes it under a fixed global budget, and invokes a single rubric-guided LLM judge to produce structured diagnostic reports; on public Magentic-One-style and tau-bench-style trajectory datasets it improves failure-localization alignment with human annotations by roughly 20–35 percentage points on Magentic-One and up to 23 percentage points on tau-retail compared with AgentRx, while reducing cost by about 6x and evaluation time by more than 8x, and it has also been deployed in the authors' enterprise agentic platform.
The work presents LiteTrajEval, a lightweight architecture for budget-bounded trajectory evaluation that derives compact domain-specific rule profiles offline, then preprocesses each trajectory online, marks heuristic failure signals, serializes it under a fixed global budget, and invokes a single rubric-guided LLM judge to produce structured diagnostic reports; on public Magentic-One-style and tau-bench-style trajectory datasets it improves failure-localization alignment with human annotations by roughly 20–35 percentage points on Magentic-One and up to 23 percentage points on tau-retail compared with AgentRx, while reducing cost by about 6x and evaluation time by more than 8x, and it has also been deployed in the authors' enterprise agentic platform.