acceptodds
Under review as a conference paper at ICLR 2027

RedTeachingBench : A Curriculum-Grounded Red Teaming Benchmark for Teachers using LLMs

Abstract

Safety benchmarks for Large Language Models (LLMs) predominantly measure the generation of hazards such as sexual content, violent expression, or hate speech. These categories are essential for safety but too generic to assess LLM reliability for high-school teachers within a given national context. We introduce RedTeachingBench, a curriculum-grounded benchmark to adversarially test small LLMs, which are particularly relevant in educational context as they comply with the requirements of national sovereignty, confidentiality, and model frugality. Our benchmark tests whether LLMs elaborate on false premises, creating a misinformation risk for students. The benchmark is built following two methodological steps: a) with the help of volunteering French teachers, we extract key knowledge in six disciplines for students between 12 and 15; b) for each of 500 knowledge items, we generate two kinds of false-premise questions through minimal semantic corruption, distinguishing plausible local errors from overtly absurd substitutions. We evaluate RedTeachingBench on the generation of flashcards. Initial tests on small LLMs show substantial failures: models achieving a mean of 86.7% safe-answer rates on a standard red-teaming benchmark as AdvBench correctly flag only 41% of our examples, and fully reject only 7.9%. However, results differ significantly with model alignment and reasoning capabilities of LLMs. It is also noteworthy that reasoning models demonstrate different performance on the two kinds of false-premise questions. These findings provide preliminary evidence that our curriculum-grounded adversarial benchmark measures a distinct robustness property against misinformation, which is relevant for real-life deployment in the educational sector.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.