An ensemble-based information-theoretic task selection framework makes LLM multi-agent communication-structure optimization beat random training in both benign and adversarial settings
Synopsis
The work proposes an ensemble-based information-theoretic task selection framework that estimates task informativeness by how much a candidate task changes the distribution over graph parameters, using ensemble Kalman inversion as a derivative-free approximation of the Bayesian update, and combines embedding-based representative selection, surrogate modeling, and batch Thompson sampling for scalability; validated across benign and adversarial settings and multiple task formats, it consistently outperforms random training with more effective task selection and greater overall cost efficiency.
Figure 1: Accuracy-cost scaling on MMLU (adversarial). More randomly selected training tasks help only modestly, while active learning achieves higher accuracy despite lower cost.
arXivInterpretation
It introduces an information-theoretic selection criterion that measures task informativeness by how much a candidate task changes the distribution over graph parameters, for communication-structure optimization in LLM-based multi-agent systems. Existing methods typically rely on randomly sampled training tasks, yet tasks differ substantially in difficulty and domain and are not equally informative for updating communication structure, making optimization unstable and highly sensitive to the particular training set; this work shifts task selection from random sampling to actively identifying the most valuable tasks. The abstract states the motivation and design and reports validation across benign and adversarial settings and multiple task formats showing consistent improvement over random training.
It uses ensemble Kalman inversion as an efficient, derivative-free approximation of the corresponding Bayesian update, making the informativeness estimator suitable for black-box and noisy multi-agent systems. By pairing a Bayesian-update approximation with an ensemble approach, it avoids differentiating through the system, fitting multi-agent environments where only inputs and outputs are observable. The abstract explicitly states the estimator is especially suitable for black-box and noisy multi-agent systems and positions it as an approximation of the Bayesian update.
It builds a compact candidate pool through embedding-based representative selection and combines informative selection with surrogate modeling and batch Thompson sampling to enhance scalability. Beyond information-theoretic selection, it adds candidate-pool compression and batch sampling so the framework can operate over larger task spaces. The abstract lists these three components as the scalability design and describes how they combine with informative selection.
Validated across benign and adversarial settings and multiple task formats, the framework consistently outperforms random training and shows greater overall cost efficiency. Compared with the random-training baseline, the framework improves both task-selection effectiveness and cost efficiency. The abstract reports consistent advantages across settings and task formats but gives no specific numbers, sample sizes, or statistical-test details.
Perspective
The framework targets communication-structure optimization in LLM-based multi-agent systems and applies to training settings where tasks differ substantially in difficulty and domain and where the system is observable only as a black box with noise; its design goal is to actively pick more informative tasks under such conditions instead of random-sampled training. The validation described in the abstract covers benign and adversarial settings and multiple task formats, indicating the scope is not limited to a single task form. For researchers and practitioners who want to lower the training cost of communication-structure optimization or obtain more robust updates under unstable task distributions, this framework offers a reusable task-selection approach.
The abstract reports no specific performance metrics, token savings, task counts, or statistical significance, so the magnitude of the cost-efficiency advantage remains to be confirmed in the full text. As an approximation of the Bayesian update, ensemble Kalman inversion's approximation error across noise levels and graph-parameter dimensions is an open question worth watching. The abstract does not separate the individual contributions of candidate-pool compression and batch Thompson sampling to the final results. The concrete construction of the adversarial settings and the range of task formats are also not detailed, and these affect how far the conclusions can be extrapolated.
