AIMPACT: An Evidence-Grounded Multi-Agent Framework for Automated Presentation Assessment
Abstract
Automated presentation assessment is essential for scalable presentation training and feedback in higher education and workplace settings. Despite recent advances in Multimodal Large Language Models (MLLMs), presentation assessment remains constrained by the direct mapping of heterogeneous multimodal evidence to final assessments, which couples evidence reliability with rubric-based judgment. To address this problem, we introduce AIMPACT, an evidence-grounded multi-agent framework for automated presentation assessment. Specifically, we first propose an Evidence-to-Judgment Multi-Agent Framework (EJMA) to organize presentation assessment into a two-level hierarchy. At the first level, multimodal observations are converted into a traceable evidence memory and coordinate seven dimension-specific specialist agents through a reliability-guided collaboration graph, forming an evidence-grounded council prior. On this basis, the second level introduces Rubric-Conditioned Multi-Expert Preference Optimization (RC-MEPO), which first learns expert corrections to the council prior and then performs on-policy preference optimization against the confidence-weighted expert score distribution, where controlled rubric perturbations further enforce assessment consistency with the provided scoring rubric. In addition, we construct a Multi-Expert Presentation Assessment Dataset (MEPA) that comprises 35 graded classroom presentations with synchronized audio, video, and slides. Five expert instructors independently assessed each presentation across eight rubric dimensions, resulting in 1,400 individual expert judgments. Extensive objective and subjective evaluations show that our method achieves the strongest presentation-level expert alignment among comparable-scale MLLMs and remains competitive with substantially larger proprietary models. Zero-shot evaluations on out-of-domain datasets and additional analyses further demonstrate its ability to generalize to unseen presentation data and different scoring standards.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.