Abductive learning of intermediate video concepts in a neuro-symbolic AI Architecture
Abstract
A network trained to detect a compositional video event must discover the intermediate concepts that define it from examples alone; when the training data is narrow, it latches onto surface patterns of the training distribution rather than the event itself, and fails on unseen scenarios. Supervising those concepts directly fixes this but demands expensive per-frame, per-region annotation. We take the neuro-symbolic route instead: domain knowledge enters the model as differentiable logic rules and auxiliary symbolic induction losses. We couple a grid-adapted VideoMamba perception backbone with Scallop, a differentiable logic programming framework, and evaluate on a purpose-built synthetic testbed of role-based agents delivering instruments on a discrete grid, in which every intermediate concept has exact per-frame ground truth. With zero intermediate labels, the model reaches Pearson correlation above on all three core channels on the training distribution; of these, only instrument possession is recovered without an appearance prior. On unseen scenario complexity its decisions far exceed purely neural baselines sharing the same backbone ( against on the extreme set), while remaining comparable or worse under pure perceptual degradation such as occlusion; models trained with less expensive supervision produce fewer false positives while recovering the same events. A component ablation separates the symbolic machinery into structural gradient routing from the logic rules and visual grounding from the auxiliary losses. Either alone yields comparable decisions, but only grounding yields interpretable intermediate concepts, and the two come apart under recoloring in both directions: only their combination transfers on the clean recolored set ( against and ), while under combined recoloring and extreme complexity all three degrade and routing alone is best ( against ). Decision accuracy is therefore a poor proxy for concept quality in neuro-symbolic systems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.