Geometry-Conditioned Adaptation for CLIP-Based Cross-Domain Facial Beauty Prediction
Abstract
Facial beauty prediction (FBP) estimates ratings from particular annotator populations, yet imaging conditions and rating protocols vary across datasets. Landmark geometry encodes facial proportions and anatomical relationships without pixel-level texture or illumination, offering a complementary structural cue that may be less sensitive to changes in image appearance. This could improve robustness to domain shifts across datasets. Source-tuned CLIP provides transferable, language-aligned representations, but directly fusing image-derived landmarks may amplify noisy or redundant features and perturb the image–text alignment used for prompt-based scoring. We propose a geometry-conditioned tangent adapter that keeps the CLIP image embedding as the prediction anchor. A separately trained landmark–contour encoder supplies structural features that condition image-dependent corrections. With both encoders frozen, the adapter projects its correction onto the tangent plane, bounds its norm, and re-normalizes the image embedding before applying the original text prompts and scoring rule. Source-training statistics calibrate adapter inputs, and zero initialization recovers the frozen scorer before adapter training. We evaluate on LiveBeauty, MEBeauty, and SCUT-FBP5500 under three within-dataset and two cross-dataset protocols without target-domain adaptation. The full model improves all three correlation metrics over the reported state-of-the-art FPEM results across all five settings. For LiveBeauty-to-SCUT-FBP5500 transfer, the full system's relative SRCC, PLCC, and KRCC gains over FPEM are , , and , respectively. In the same setting, PLCC rises from for our source-tuned CLIP base to for the full model ( relative to that base). These results support bounded geometry-conditioned adaptation as a useful design for this CLIP-based FBP system without target-domain fine-tuning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.