Skip to main content
Back to timeline
arXivSource publication:

EvoCast separates cognition from authority to evolve forecasting architectures autonomously, reaching the lowest MSE on three real forecasting cases at lower agent-side cost

Related research and updates

Synopsis

EvoCast is a fully autonomous research-agent system for iterative time-series forecasting architecture evolution: it first establishes and diagnoses a task-specific baseline, then generates evidence-grounded research directions from dataset characteristics, diagnostic results, prior rounds, and failure records, with LLM agents handling hypothesis generation and code implementation while deterministic program authorities control source-edit boundaries, canonical evaluation, and model promotion; it reaches 96.7% success on 30 repository-level implementation tasks and the lowest MSE on three real forecasting cases (Air Quality, Melbourne Pedestrian T33, Steel Industry) while using fewer tokens and lower API cost than R&D-Agent.

Source-provided article image: EvoCast: Reliable Autonomous Research Agents for Iterative Forecasting Architecture Evolution
Figure 1 ·

Figure 1. EvoCast in the landscape of automated forecasting model development. Compared with fixed-space AutoML/NAS and open-ended LLM research agents, EvoCast combines evidence construction, bounded cognitive-agent exploration, and program-authority verification in a fixed-protocol loop for reliable architecture evolution. A three-part comparison of fixed-space AutoML and neural architecture search, open-ended LLM research agents, and EvoCast. EvoCast links task evidence, bounded cognitive-agent exploration, program-authority verification, canonical evaluation, and evidence-state updates in a closed loop.

arXiv

Interpretation

EvoCast formulates autonomous forecasting architecture evolution as a fixed-protocol research loop spanning baseline establishment, mechanism diagnosis, bounded implementation, canonical evaluation, promotion, and evidence update in heterogeneous forecasting repositories. Unlike AutoML/NAS constrained by predefined search spaces, and unlike general LLM research agents that handle research generation, code execution, and result adjudication inside one open-ended agent, this work binds the whole process to a frozen task protocol and a single comparison surface. The paper supports this loop with explicit formalization of protocol freezing, baseline selection, mechanism ablation, candidate execution, and promotion rules, executed on the unified TFB benchmark path for baselines, ablations, and candidates.

It introduces cognition-authority separation: LLM agents handle open-ended hypothesis generation, source implementation, and report narration, while deterministic program authorities control source validity, canonical evaluation, metric comparison, model promotion, and state updates. This design explicitly separates research-intent generation from experimental adjudication, so agent self-report cannot directly become an experimental conclusion; only candidates that pass boundary validation, produce valid canonical metrics, and satisfy the fixed multi-seed comparison rule can update the current best model. The paper defines three permission levels, boundary enforcement, and two-level promotion semantics (evidence recording versus incumbent replacement), and the RQ3 A4 ablation shows that removing authority control raises runtime, tokens, and cost to 10.80 hours, 25,321,252 tokens, and 7.8792 USD while the retained final artifact stays at the baseline level of 0.185001.

It builds an evidence-grounded execution framework that turns research directions into provenance-tracked source candidates, uses repair and failure memory to reduce ineffective iteration, and accumulates scientific and engineering evidence across rounds. Compared with agent workflows that rely mainly on execution feedback or runtime repair, EvoCast writes ablation outcomes, accepted improvements, valid negatives, unstable gains, and engineering failures separately into the evidence state so later planning reads the full research trajectory. The RQ3 A5 ablation (freezing iterative evidence update) degrades final MSE to 0.178489 with 15/20 valid rounds, versus 0.162579 and 18/20 for the full system, indicating a role for cross-round evidence write-back in sustained progress.

On repository-level implementation and end-to-end forecasting research tasks, EvoCast shows higher implementation reliability and lower agent-side cost, and produces task-specific architectures. On 30 tasks built from 10 forecasting backbones and three idea-complexity levels, EvoCast reaches 96.7% success (29/30), 100% fidelity, and 95.7% repair success (22/23), with average 27,618 tokens, 0.0137 USD, and 153 seconds; on the three real cases it attains the lowest MSE and the lowest MAE on two of them. Results come from comparison against Direct Independent Edit, mini-SWE-agent, AIDE, R&D-Agent, and AutoResearch under the same repository snapshot, idea text, and editable boundary, plus final-deliverable comparison on three public datasets under one unified TFB evaluator.

Perspective

The results apply to repository-level architecture evolution under a fixed task protocol, starting from forecasting models in the TFB repository family, and suit researchers and engineering teams who want to obtain task-specific architectures automatically while retaining auditable evidence. The output bundle includes the validated current-best model, reproducible source and configuration, resource usage, and an HTML report, so it can be used directly for reproduction, ablation inspection, and result export; Build mode further supports fast end-to-end reliability checks before longer formal runs. The method's value lies in continuously converting budget into task-specific modules that remain on the original backbone's live path rather than replacing it with a more conservative general solution.

Open questions remain: whether the conclusions from three real cases and a single ablation task (Melbourne Pedestrian T33) hold on more datasets, more repositories, and longer research budgets; whether the relative contribution of each authority-separation component is stable across backbones and task modes; and how the balance between boundary constraints and repair budget affects final architecture quality when candidate implementations require larger cross-file restructuring. In addition, the reported resource comparison shows EvoCast uses more tokens and API cost than AIDE on Air Quality and Steel Industry, with relative wall-clock time varying by task, so the conditions under which its cost advantage applies deserve re-checking on a reader's own tasks.

Sources