ShutterMuse: Capture-Time Photography Guidance with MLLMs
Abstract
Computational photography has largely focused on post-capture optimization, while many photographic failures could be prevented at capture time. Capture-time assistance requires more than aesthetic scoring or crop prediction: it must determine whether the current view should be preserved, refined, or abandoned, and guide how a subject should pose within the scene. We formulate this problem as capture-time photography guidance and introduce ShutterBench, a benchmark evaluating both sides of the camera. Photographer-side guidance evaluates composition decisions (keep/refine/reject) and spatial refinement, while subject-side guidance evaluates scene-conditioned pose recommendation. Our evaluation reveals complementary limitations: general-purpose MLLMs make reasonable composition decisions but struggle with precise spatial refinement, whereas specialized cropping models localize effective crops but cannot reason over the broader decision space; MLLMs also generate valid poses, but their aesthetic quality remains limited. To enable unified learning, we construct ShutterData, comprising 130K samples with textual rationales and structured composition and pose supervision, and develop ShutterMuse, a unified MLLM trained with supervised and reinforcement fine- tuning. ShutterMuse achieves the strongest overall photographer-side performance on ShutterBench and competitive subject-side performance at substantially lower inference cost, demonstrating its potential to shift computational photography from post-hoc optimization toward actionable assistance before capture.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.