MEA uses reward-driven multi-agent optimization to raise explanation faithfulness by up to 34%
Related research and updatesSynopsis
The work presents MEA, a multi-agent framework in which a Proposer agent selects and configures explanation tools based on the question and modality while an Actor agent is optimized end-to-end against faithfulness to turn outputs into natural-language explanations; the authors also introduce question types spanning feature attribution, counterfactual reasoning, and spurious feature detection, each paired with a perturbation-based faithfulness metric, find that frontier LLMs systematically produce unfaithful explanations, and report that MEA outperforms post hoc explainers, agentic, and closed-source baselines across six datasets, with faithfulness gains of +28% (tabular), +21% (text), and +34% (vision) over the untrained backbone.
Figure 1: Overview of Mea . Given a user question about model behavior, the Proposer agent reasons over the model prediction and input modality to select an explanation strategy, choosing from a multimodal (tabular, text, and vision) toolkit and agent reasoning tasks (grounding, reasoning, and comparison), the Actor agent then executes the strategy, processes raw explanation outputs, and produces a structured natural-language explanation. The explanation is evaluated via input perturbation against faithfulness metrics. In training, both agents are jointly optimized end-to-end via GRPO [ 14 ] , using a reward combining faithfulness and tool penalties.
arXivInterpretation
MEA assigns explanation tool selection and configuration to a Proposer agent and explanation generation to an Actor agent optimized end-to-end against faithfulness, producing natural-language explanations across tabular, text, and vision modalities. Whereas post hoc explanation methods require users to navigate high-dimensional outputs and synthesize evidence across tools themselves, MEA moves that knowledge burden into a multi-agent pipeline. The abstract states the framework operates across three modalities and is compared with post hoc explainers, agentic, and closed-source baselines on six datasets; implementation details are not expanded in the abstract.
The authors introduce diverse question types covering feature attribution, counterfactual reasoning, and spurious feature detection, each paired with a perturbation-based faithfulness metric. This provides a quantifiable evaluation dimension for explanation quality rather than relying on human judgment or a single tool's output. The abstract explicitly lists the pairing of question types with perturbation-based faithfulness metrics, but does not give the metric's computation or values.
The authors report that frontier LLMs systematically produce unfaithful explanations. This observation locates the problem in explanation faithfulness rather than fluency, indicating that direct language-model generation of explanations is not reliable on its own. The conclusion comes from an overall statement in the abstract; the specific model list, evaluation scale, and statistical details are not given in the abstract.
After optimizing against faithfulness rewards augmented with a modality-adaptive penalty, MEA consistently outperforms post hoc explainers, agentic, and closed-source baselines across six datasets, with faithfulness gains of +28% (tabular), +21% (text), and +34% (vision) over the untrained backbone. The gains come from reward-driven end-to-end optimization rather than a larger model or more tools, suggesting the optimization objective itself is key to improving faithfulness. The abstract reports consistent comparisons across six datasets and three-modality gain figures, but provides no confidence intervals, significance tests, or per-dataset breakdown.
Perspective
The work targets domain practitioners who need to understand model behavior but lack expertise in explanation tools, applies to tabular, text, and vision modalities, and evaluates on question types including feature attribution, counterfactual reasoning, and spurious feature detection. Its value lies in delegating explanation tool selection, configuration, and natural-language synthesis to an agent pipeline, making faithfulness an optimizable objective rather than a post hoc check; for researchers and engineering teams seeking scalable explanation interfaces, it offers a path to reward-driven explanation generation.
The abstract does not specify how the perturbation-based faithfulness metric is computed, the composition of the six datasets, baseline configurations, or the statistical significance of the gains, making it hard to judge stability across task distributions. The model range and evaluation scale behind the observation that frontier LLMs systematically produce unfaithful explanations also need confirmation in the main text. In addition, the concrete form and mechanism of the modality-adaptive penalty are not expanded in the abstract, and its contribution to cross-modality generalization remains to be examined.
