Skip to main content
Back to timeline
arXivSource publication:

CERA-MoA: A Mixture-of-Agents Framework Where Routing Mechanisms and Continually Learning Agents Co-Evolve

Synopsis

This work introduces CERA-MoA, an iterative reinforcement learning framework in which a dynamic router and multiple independent agent policies co-evolve: the router uses a predictive familiarity estimator built on mid-layer hidden states to assess each agent's semantic competence before generation, applies cumulative-threshold adaptive routing to activate a minimal agent subset, and proactively allocates targeted training samples based on agents' evolving competence, outperforming static-agent routing and fix-workflow fine-tuning baselines across diverse domains.

Source-provided article image: CERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents
Figure 1 ·

Figure 1: Overview of CERA-MoA.

arXiv

Interpretation

It proposes a sample-level Mixture-of-Agents framework that enables continual reinforcement learning of agents through adaptive query routing within one closed loop. Prior orchestration optimization assumes fixed-capability candidates, while agent fine-tuning relies on manually partitioned or uniform datasets, leaving the two decoupled; this work embeds iterative routing directly into the training loop and allocates semantic queries based on evolving competence. The paper presents the closed-loop design for training and inference phases and reports outperforming static-agent routing and fix-workflow fine-tuning baselines across diverse domains.

It designs a predictive familiarity estimator that uses mid-layer hidden states to evaluate agent-query semantic compatibility before generation. Unlike approaches relying on external evaluators or full rollout generation, it concatenates hidden states at layers floor(L/2) and floor(3L/4), projects them through a trainable predictor head and a frozen target head, and derives familiarity via normalized Euclidean distance and exponential decay, avoiding full-generation overhead. The paper specifies the estimator's construction and equations and notes that the frozen target head merely provides a stable reference point rather than encoding a semantic prototype or competence label.

It introduces cumulative-threshold adaptive routing that activates a minimal agent subset to balance performance and computational cost. Instead of fixed top-k allocation, it ranks agents by an exploration-augmented routing score and activates the smallest prefix whose cumulative raw familiarity reaches threshold tau, falling back to the full population when the total stays below tau. The paper provides equations for the ranking score and activation criterion and states that exploration bonuses and historical entropy are used during training and set to zero at inference.

It promotes capability differentiation through competence-aware sample allocation, moving agents from homogeneous generalists toward domain specialists. Agents that solve queries in a specific semantic space gain higher familiarity and become more likely to receive semantically similar queries in the future, forming a learning loop that naturally induces domain specialization. The paper adopts a zero-shot, role-agnostic prompt design and states that the emergent capability differentiation in its empirical analysis is driven by familiarity routing and closed-loop reinforcement learning rather than human-engineered role bias.

Perspective

The framework targets complex reasoning tasks requiring multi-agent collaboration and is configured for shared-backbone (N=4) and heterogeneous-backbone (N=3) settings, with zero-shot, role-agnostic prompts, an 800-token input limit, a 1000-token generation limit, and built-in reasoning modes disabled. It lets researchers jointly optimize routing and agent policies within the training loop and lets the router activate a minimal subset at inference by familiarity threshold; deterministic tasks use familiarity-weighted majority voting, while open-ended tasks output the response of the most familiar agent.

The loaded text does not include specific experimental result tables or numbers, so the magnitude of performance gains across domains, the quantitative evidence for capability differentiation, and the size of efficiency benefits still need to be confirmed from the original figures and tables; in addition, the stability of the familiarity estimator when the backbone changes or domains shift substantially, and the sensitivity of threshold tau and separation margin m, are open questions worth watching.

Sources