LiveMACEBench evaluates five frontier LLM agents over 30 days of live markets and finds return rankings diverge sharply from capability metrics
Synopsis
The work introduces LiveMACEBench, a process-aware benchmark that uses live cryptocurrency and U.S. equity markets as a naturally evolving testbed in which five frontier LLMs run continuous paper-trading trajectories under matched Base, Tool Use, Persistent Memory, Rule Following, and Multi-Agent Collaboration configurations, evaluated through both realized outcomes and trace-derived diagnostics; across roughly 30 days it finds a pronounced outcome–capability gap, with return rankings fluctuating and often diverging from capability-specific measurements while similar outcomes arise from markedly different patterns of mechanism use.
Figure 1 : LiveMACEBench overview. Agents interact with a shared, continuously evolving live-market environment along persistent trajectories. Evaluation combines realized outcomes with mechanism-specific process diagnostics to distinguish task performance from how effectively agents use each mechanism.
arXivInterpretation
It formulates process-aware agent evaluation that separates mechanism access, effective mechanism use, and realized task benefit, combining task outcomes, controlled mechanism effects, and trace-derived diagnostics. Prior agent benchmarks largely rank final task success or returns; this work shifts the object of evaluation to the mechanisms that produce outcomes and builds matched configurations on a common ReAct scaffold so capability differences can be localized rather than inferred from aggregate performance. The design specifies five controlled configurations with complete decision-trace recording; empirically it collects traces from five frontier LLMs over roughly 30 days of live markets and contrasts return-ranking reversals with capability-ranking reversals (21 versus 0 for tool use, 19 versus 5 for memory, 19 versus 8 for rule following, 29 versus 2 for collaboration).
Within tool use, information coverage is separable from invocation efficiency, and tool discovery is separable from tool execution: all four models exceed 0.8 information coverage while invocation efficiency is consistently the weakest judge-rated dimension, Qwen3-Max has the lowest routing quality yet a near-perfect valid call rate, and Grok-4.20 has higher routing quality but a valid call rate of 0.666. It decomposes tool use into discovery, invocation, and evidence use, and pairs LLM-as-a-judge dimensions with objective trace metrics so that finding a tool and using it well are no longer conflated. Objective metrics (routing quality, valid call rate, hallucination-free rate) are computed by replaying recorded tool traces, alongside multi-judge scoring with reported cross-judge robustness showing stable model ordering but judge-dependent absolute scores.
Persistent memory is a multi-stage capability in which storage volume, content quality, and effective use are distinct, and memory yields consistent downside-risk improvement rather than uniform return gains: all five memory-augmented accounts reduce tail loss relative to matched memory-free counterparts, while return effects differ across models. It extends memory evaluation from a single storage or retrieval measure to a chain of construction, retrieval, adoption, and downstream risk effect, using matched memory-free accounts as the comparison condition. It reports per-model retrieval and storage counts, content-quality and usage-quality scores, and changes in maximum drawdown and tail loss relative to matched memory-free accounts; case studies further show memory protecting capital in range-bound regimes while capping upside in a trending regime.
Rule following requires two complementary competencies, execution-level compliance and auditable rule reasoning, which can diverge; likewise, in multi-agent collaboration, role participation and evidence integration do not move monotonically with realized return. It measures execution compliance with a programmatic checker, rule understanding and conflict handling with an independent LLM auditor, and observable collaboration with role diversity and evidence integration, separating what was done from what can be explained and used. It reports hard-rule pass rate, rule satisfaction score, and rule auditability score with cross-auditor consistency analysis, and case studies illustrating silent compliance, hallucinated violation repair, boundary targeting, and auditable violation repair.
Perspective
The benchmark targets persistent LLM agents operating in continuously evolving environments, and is meant for evaluation settings where mechanism access, effective mechanism use, and realized benefit need to be distinguished, such as long-horizon decision making, tool augmentation, memory augmentation, policy-constrained behavior, and multi-agent collaboration. Its instantiation uses live cryptocurrency and U.S. equity paper trading with four-hour decision rounds and finer-grained account-equity checkpoints, making it most directly usable by researchers and engineering teams who want to compare backbones or mechanism configurations under shared external conditions. The findings apply to the observed roughly 30-day window and its market conditions, and the authors note that extending process-aware evaluation to longer horizons, broader regimes, and other continuously evolving environments is a natural next step.
Several open questions remain for a careful reader: whether the relationship between capability metrics and returns holds over longer horizons and different market regimes needs more windows to establish; some judge dimensions such as invocation efficiency and evidence faithfulness show lower agreement across judge models, and absolute scores depend on judge strictness; role-set conventions and the role-dominance threshold shift absolute collaboration scores even though model ordering is preserved; and individual runs were excluded because of an environment consistency defect or an incomplete trajectory, with the scope of that exclusion described in the text. In addition, although this is the full text, some tables and figures are missing numeric values in the parsed version, so specific multi-agent return differentials and some rule-level figures cannot be checked item by item here.
