acceptodds
Under review as a conference paper at ICLR 2027

Species-Aware Morphometric Priors for Avian 3D Reconstruction

Abstract

Reconstructing a bird in 3D from one photograph remains hard. There are 11,000 species of extreme variety, almost no 3D shape data, and no absolute scale in the image. Annotations mark one point per part, which says where a part is but not how long it is. Current methods regress the parameters of a bird model built from a few species, and they recover pose and silhouette well. Yet the part proportions they produce barely change from one species to another and do not match those measured for the species. Our starting observation is a difference between birds and humans. Humans remain one species however skin, height or limb length varies. Birds share the same joints, yet how far the bill, tarsus, tail and wing extend from those joints differs by species and is likely what tells one species from another. Ornithology measures exactly these lengths, and we bring them into 3D reconstruction from AVONET, an open database of more than 90,000 measured individuals across 11,000 species. Since a photograph has no scale, we follow the AVONET definitions and express each length as a ratio to wing length. Since these ratios vary little within a species, the species-average ratios serve as a species-aware morphometric prior. A morphometric pathway realizes it without disturbing the pretrained body model. A proportion head infers per-part length factors from the features of a pretrained transformer reconstructor, and an axial deformation operator inside the bird model applies them before skinning. Because the body model is left intact, pose and camera adapt to the new part lengths and species morphology is added without relearning the shape space. No photograph carries a measurement: for real images training uses species averages alone, and inference uses one photograph. On 49 held-out species the median error of the tail, bill and tarsus proportions against the measured species values falls from 21–31% without the prior to about 10%. On 237 species never seen in training it falls from 25–61% to 11–18%, while 2D keypoint accuracy improves and the 3D error is preserved. Long-legged species such as herons, cranes and storks, absent from the training photographs, remain difficult, and extending the prior to them is the next task.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.