RESEARCH · RESEARCH · #1508
MEA: Reward-driven multi-agent system for faithful model explanations (arXiv:2610.02480v1)
This paper introduces MEA, a two-agent framework (Proposer and Actor) that selects and configures explanation tools and is optimized end-to-end against perturbation-based faithfulness rewards, producing natural-language explanations across tabular, text, and vision modalities. Evaluated on six datasets, MEA reportedly outperforms post-hoc explainers, agentic, and closed-source baselines and yields faithfulness gains of roughly +28% (tabular), +21% (text), and +34% (vision) over the untrained backbone, while the authors find that frontier LLMs often produce unfaithful explanations.
KEY POINTS
- This paper introduces MEA, a two-agent framework (Proposer and Actor) that selects and configures explanation tools and is optimized end-to-end against perturbation-based faithfulness rewards, producing natural-language explanations across tabular, text, and vision modalities.
- Evaluated on six datasets, MEA reportedly outperforms post-hoc explainers, agentic, and closed-source baselines and yields faithfulness gains of roughly +28% (tabular), +21% (text), and +34% (vision) over the untrained backbone, while the authors find that frontier LLMs often produce unfaithful explanations.
- Shows that optimizing agentic explainers directly for perturbation-based faithfulness can materially improve explanation fidelity across modalities and exposes systematic unfaithfulness in frontier LLM explanations.
WHY IT MATTERS
Shows that optimizing agentic explainers directly for perturbation-based faithfulness can materially improve explanation fidelity across modalities and exposes systematic unfaithfulness in frontier LLM explanations.