Solver-JEPA: A Joint-Embedding Predictive Architecture for Visual Problem Solving
Abstract
Visual problem solving requires transforming visually presented problems into structured solutions under global constraints and objectives. While methods learn to generate solution images or reasoning trajectories, visual problem solving still lacks a representation-learning objective designed to organize the information required for solution construction. Such learning is challenging because the visually observed structure and the underlying decision structure are not equivalent: small local differences can induce discontinuous changes in global feasibility or optimality, while a final solution image reveals only one successful endpoint and leaves the consequences and interactions of alternative decisions unspecified. We introduce **Solver-JEPA**, a solver-supervised joint-embedding predictive architecture that turns each problem into a family of decision-conditioned prediction tasks. During training, partial commitments specify tentative local decisions, and task solvers provide targets describing their consequences for problem structure, global feasibility, attainable solution quality, and compatible complete solutions. Predicting these targets in representation space encourages the visual encoder to preserve solution-relevant local-to-global dependencies. A separately trained decoder reuses the frozen representation to construct complete solution images without solver access at inference. Across four visual problem solving tasks, Solver-JEPA improves strict solution accuracy by up to **11.6** percentage points and provides **15**–**150 times** faster inference than generation-based baselines. Controlled ablations show that joint-embedding alignment and composite-commitment supervision improve the readability of complete solutions from the learned representations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.