RuleRouter: Sparse, Traceable Natural-Language Interventions for Frozen Vision-Language Models
Abstract
Natural-language rules can target structured errors in frozen vision-language models, yet unconditional application often overturns correct predictions. We introduce RuleRouter, a sparse and traceable intervention layer that derives contrastive rules from development-only confusion evidence and applies a proposed correction only when at least four of five benefit–harm router replicas agree that predicted benefit exceeds harm. The rules, routers, and operating points are frozen before their corresponding confirmation evaluations. Relative to a supervised logit-adapted System-1, RuleRouter improves Macro-F1 on EuroSAT-RGB by 5.53 percentage points (95% paired-bootstrap CI [3.90, 7.06]) at 16.8% intervention coverage. On an independent AID-30 final-confirm set, a separately frozen pipeline improves its prespecified reporting comparator by 3.17 points (CI [2.26, 4.15]), while changing 7.83% of its predictions. So2Sat serves as a low-base-performance stress test and shows a smaller 0.46-point gain. Across these evaluations, selective routing produces more corrections than harms and avoids the degradation caused by unconditional rule fusion. Equal-label visual probes remain more accurate overall, while matched semantic controls favor confusion-driven rules on EuroSAT but provide no corresponding evidence on So2Sat. RuleRouter therefore addresses settings that require sparse changes to a frozen predictor, explicit coverage and harm accounting, and an inspectable textual record for each intervention.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.