JOVE jointly assigns executor LLMs and paid verification, cutting average cost and latency by at least 3.17x while staying competitive on accuracy across four reasoning benchmarks
Related research and updatesSynopsis
JOVE is an online framework that decomposes complex reasoning queries into directed acyclic task graphs and distributes them across heterogeneous LLMs while jointly deciding which intermediate outputs to send for paid verification; it makes execution and verification decisions by solving a sequence of per-query mixed-integer linear programs, updates task-dependent LLM quality estimates online from verification feedback, and adds an information-gain bonus that folds the value of learning into allocation, achieving sublinear quality-learning regret under a long-term budget and a per-query latency constraint, with competitive accuracy against standard inference baselines and at least 3.17x lower average cost and latency across four reasoning benchmarks.
Figure 1: Overview of JOVE. A planner decomposes each query into a task graph and identifies candidate executor LLMs for each node. JOVE jointly assigns executors and selects outputs for verification, subject to a long-term execution/verification budget and latency constraints. Verification proceeds asynchronously without delaying the current response, improving model-quality estimates.
arXivInterpretation
It proposes JOVE, an online framework that jointly couples executor-LLM assignment with the selection of intermediate outputs for paid verification, where verification runs asynchronously and is used to improve future allocations. Prior work that distributes task graphs across heterogeneous LLMs typically schedules only the execution side, whereas this work weighs spending on execution now against spending on verification for later learning within a single decision. The paper presents this contribution through the framework description and algorithm design, stating that verification runs asynchronously and that its feedback flows back into future allocations; the abstract gives no concrete figures for verification budget or verification frequency.
It formulates the trade-off as a stochastic optimization problem under a long-term budget and a per-query latency constraint, with initially unknown service quality, invocation costs, and execution times, and solves it as a sequence of per-query mixed-integer linear programs. It casts resource-aware LLM task-graph scheduling as a solvable sequence of MILPs rather than a heuristic rule, so budget and latency constraints can be expressed explicitly. The abstract states that JOVE makes execution and verification decisions by solving a sequence of per-query mixed-integer linear programs, and specifies the setting of stochastic, initially unknown service quality, invocation costs, and execution times.
It updates task-dependent LLM quality estimates online from verification feedback and adds an information-gain bonus that incorporates the value of learning into allocation decisions. Exploration enters the allocation objective directly as an information-gain term, making the trade-off between exploiting the currently best executor and acquiring information for the future explicit. The abstract states that online learning updates task-dependent quality estimates, that an information-gain bonus is incorporated into allocation decisions, and that sublinear quality-learning regret is established under a natural set of assumptions.
Across four reasoning benchmarks, JOVE achieves competitive accuracy against standard inference baselines while reducing average cost and latency by at least 3.17 times. It evaluates resource metrics (cost and latency) alongside accuracy rather than accuracy alone, and reports a quantified resource saving. The abstract reports results on four reasoning benchmarks and gives the at-least-3.17x cost and latency reduction; it does not list the benchmark names, baseline configuration details, or per-item numbers.
Perspective
The work targets online serving scenarios that decompose complex reasoning queries into directed acyclic task graphs and execute them across heterogeneous LLMs, in deployment settings with a long-term budget cap and a per-query latency constraint where each LLM's service quality, invocation cost, and execution time for a given subtask are initially unknown. It lets a system allocate resources between execution and paid verification and lets verification feedback continuously improve later allocations, so it is directly relevant to engineering teams that must control inference cost and response time while relying on smaller models for complex tasks. The theoretical result is sublinear quality-learning regret established under a natural set of assumptions, and the empirical conclusions come from comparisons against standard inference baselines on four reasoning benchmarks.
The abstract does not name the four reasoning benchmarks, does not specify the configuration of the standard inference baselines, and does not give per-item accuracy, cost, or latency numbers, so the conditions under which the at-least-3.17x reduction is obtained, in terms of benchmark, budget, and latency constraint, still need to be confirmed in the paper. The content of the natural set of assumptions behind the sublinear quality-learning regret, the magnitude of verification fees, and the setting of verification frequency are not expanded in the abstract. In addition, how verification errors would affect allocation decisions, and where the trade-off boundary between the long-term budget and the per-query latency constraint lies, are directions a reader may continue to watch.
