AnyPoint: Dynamic Evaluation for Point-Cloud Multimodal Large Language Models
Abstract
Point-cloud multimodal large language models (MLLMs) are advancing rapidly in 3D scene understanding, yet their evaluation relies almost exclusively on static benchmarks that quickly saturate as models improve and fail to locate individual model weaknesses. Transitioning to dynamic evaluation requires an expansive pool of physically plausible 3D scenes with verified ground-truth answers, a dual challenge that neither manual annotation nor pure language model generation can satisfy. To resolve this dilemma, we present AnyPoint, a dynamic evaluation framework that couples programmatic 3D scene generation with adaptive failure discovery for point-cloud MLLMs. AnyPoint decouples scene creation into semantic layout generation via a language model and spatial coordinate solving via a deterministic constraint solver, producing physically valid 3D scenes alongside grounded 3D scene graphs. To avert saturation, evaluation tasks are derived compositionally as programs over each graph, allowing ground-truth answers to be computed directly from the geometry without human annotation. Within this generated task space, a budget-aware dynamic evaluation algorithm adaptively probes each model's failure boundary by balancing error exploitation and novelty exploration. Across six representative point-cloud MLLMs, AnyPoint raises the failure discovery rate within a 1,000-query budget from 65.80% under uniform sampling to 83.73%. The recovered failure profiles reveal distinct model-specific weaknesses alongside shared capability bottlenecks that single aggregate scores obscure. Finally, we demonstrate that discovered failures can drive targeted model improvement, while preserving task diversity is essential for overall generalization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.