Public articles linked to the same research event.
arXiv The work presents MEA, a multi-agent framework in which a Proposer agent selects and configures explanation tools based on the question and modality while an Actor agent is optimized end-to-end against faithfulness to turn outputs into natural-language explanations; the authors also introduce question types spanning feature attribution, counterfactual reasoning, and spurious feature detection, each paired with a perturbation-based faithfulness metric, find that frontier LLMs systematically produce unfaithful explanations, and report that MEA outperforms post hoc explainers, agentic, and closed-source baselines across six datasets, with faithfulness gains of +28% (tabular), +21% (text), and +34% (vision) over the untrained backbone.
The work presents MEA, a multi-agent framework in which a Proposer agent selects and configures explanation tools based on the question and modality while an Actor agent is optimized end-to-end against faithfulness to turn outputs into natural-language explanations; the authors also introduce question types spanning feature attribution, counterfactual reasoning, and spurious feature detection, each paired with a perturbation-based faithfulness metric, find that frontier LLMs systematically produce unfaithful explanations, and report that MEA outperforms post hoc explainers, agentic, and closed-source baselines across six datasets, with faithfulness gains of +28% (tabular), +21% (text), and +34% (vision) over the untrained backbone.
The work presents MEA, a multi-agent framework in which a Proposer agent selects and configures explanation tools based on the question and modality while an Actor agent is optimized end-to-end against faithfulness to turn outputs into natural-language explanations; the authors also introduce question types spanning feature attribution, counterfactual reasoning, and spurious feature detection, each paired with a perturbation-based faithfulness metric, find that frontier LLMs systematically produce unfaithful explanations, and report that MEA outperforms post hoc explainers, agentic, and closed-source baselines across six datasets, with faithfulness gains of +28% (tabular), +21% (text), and +34% (vision) over the untrained backbone.
The work presents MEA, a multi-agent framework in which a Proposer agent selects and configures explanation tools based on the question and modality while an Actor agent is optimized end-to-end against faithfulness to turn outputs into natural-language explanations; the authors also introduce question types spanning feature attribution, counterfactual reasoning, and spurious feature detection, each paired with a perturbation-based faithfulness metric, find that frontier LLMs systematically produce unfaithful explanations, and report that MEA outperforms post hoc explainers, agentic, and closed-source baselines across six datasets, with faithfulness gains of +28% (tabular), +21% (text), and +34% (vision) over the untrained backbone.