SAHyper: Breaking the Rank Bottleneck and Rethinking Pretrained Weight Inheritance in Graph Hypernetworks
Abstract
Graph HyperNetworks (GHNs) generate parameters from computational graphs, enabling a single generator to serve models across scales and tasks. Recent methods incorporate pretrained knowledge to improve the quality of generated parameters. However, existing decoders limit the rank of generated weights, and increasing their rank is costly. Existing weight inheritance also treats related pretrained attention matrices independently. We introduce SAHyper, a Structure-Aware Hypernetwork that improves both direct generation and pretrained weight inheritance. Its Multi Component Decoder sums low-rank components to enable higher-rank weight generation with a compact hypernetwork. Structure-Aware Spectral Inheritance combines pretrained singular values with generated orthogonal bases and coordinates related bases within attention blocks. On ImageNet-1K, SAHyper generates a ViT-Base model that achieves 63.86% Top-1 accuracy without further training, surpassing the SOTA method by 10.5 points. It also outperforms the SOTA method on Visual Domain Decathlon and sets of 14 and 20 image classification tasks, with gains maintained after further training and on unseen tasks. SAHyper also outperforms the SOTA method in ResNet and LLaMA experiments.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.