GEB-Bench: Abstract Structures Told in Many Voices
Abstract
Analogy allows models to apply relational knowledge from one situation to another. Benchmarks of equivalent representations test different ways of expressing the same underlying problem, whereas analogy often links situations involving different objects and content. To evaluate the analogy ability of models, we introduce GEB-Bench, a benchmark centered on cross-voice motif matching. A voice presents a motif as a scene, story, mathematical statement, or programmatic skeleton. The first task is identification across these voices. The second is matching scenes to stories or mathematical statements, and stories to mathematical statements, by their shared motif. A complete construction pipeline generates items in these voices from the relation specified for each motif, checks the order and repetition of story units, and assembles paired questions. We evaluate current models on our benchmark. Results show that matching performance depends on the target voice. The paired scene-to-story study reveals that (1) evaluated models still make matching errors when they separately identify the source and every candidate correctly, and (2) identification improves more consistently than scene-to-story matching across the evaluated Qwen model sizes. These findings show that successful motif identification does not ensure reliable matching across voices.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.