Joint Learning For Robust And Efficient Neural Grammar Induction
Abstract
Grammar induction is the process of recovering the latent hierarchical structure of sentences without supervision, a notoriously unstable process, with different seeds converging to different grammars. While recent work on grounding probabilistic context-free grammar induction using vision as a fixed auxiliary signal showed modest improvement, it did not resolve the instability problem. Rather than treating vision as a fixed signal, we present a joint-learning approach that learns syntax and a shared image–text semantic space simultaneously, reducing variance across seeds while increasing the accuracy of the induced grammar. We test the robustness of our approach along three dimensions: across languages (English, French, and Chinese), across datasets (from Abstract Scenes to MS-COCO), and across tasks on a downstream compositional benchmark (COCO-Order task of ARO). The advantages of joint learning replicate across all three dimensions, where we highlight the importance of selecting the appropriate data to avoid collapse onto trivial grammars, introducing a harder negative-sampling regime. Furthermore, using the induced grammar model as an image–text encoder and scorer improves over CLIP on the ARO word-order task by roughly points. Finally, we demonstrate that the quality of a model's induced grammar predicts its word-order accuracy, evidence that more linguistically aligned representations generalise better.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.