MESA-Bench: A Multimodal Evolving Student Attack Benchmark for Educational LLM Tutors
Abstract
Multimodal large language models (MLLMs) are increasingly deployed as tutor models to provide problem-solving guidance, interpret visual materials, and respond to students' states across multi-turn conversations. However, existing educational tutor safety evaluations largely rely on static samples or fixed student strategies, making it difficult to capture dynamic risks in which student attacks adapt to a tutor's prior responses. We propose MESA-Bench, a Multimodal Evolving Student-Attack benchmark for tutor integrity. MESA-Bench organizes evaluation as model- and objective-specific evolution tasks, and operationalizes tutor integrity along three dimensions: pedagogical integrity (PI), protective integrity (PrI), and epistemic integrity (EI), corresponding to answer leakage, neglect of welfare or crisis cues, and judgment retreat under unreliable evidence or social pressure. To cover the nonlinear and multimodal nature of educational attacks, MESA-Bench combines education-theory-inspired attack primitives, paired multimodal mathematics cases, and feedback-driven strategy-graph evolution. Across eight MLLMs and three integrity objectives, evolved graphs improve attack success rate in 18 of 24 model-objective runs and leave the remaining six unchanged; overall ASR increases from 48.92% to 52.92%, with the largest gain on EI. Further analyses show that graph expansion is not monotonically equivalent to stronger attacks, and that input-format risk profiles vary substantially by model and objective rather than exhibiting a uniform modality effect. Code, data, and generated evaluation cases are available at https://anonymous.4open.science/r/mesa-bench-review-3D1B/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.