Hyper-Representations as Generative Models: Sampling Unseen Neural Network Weights
Konstantin SchürholtBoris KnyazevXavier Giró-i-NietoDamian Borth
Introduces a generative approach using layer-wise loss normalization to sample diverse, high-performing neural network weights directly from model zoos for effective initialization, ensembling, and transfer learning without requiring underlying training data.
Organizations increasingly produce and store massive collections of trained machine learning models across online hubs, yet repurposing the underlying knowledge captured in these model populations remains a major challenge. Prior techniques either require direct access to the original proprietary training data or rely on representations that produce dysfunctional models when generating new network weights. The article addresses this gap by developing a generative framework that directly learns from populations of neural network weights—known as model zoos—without needing underlying data samples or class labels.
The main objective of the article is to demonstrate that an autoencoder framework can learn a compressed, regularized representation of neural network weights and reliably generate new, functional, and diverse neural network weights in a single computational pass. To evaluate this, the authors introduce a novel layer-wise loss normalization technique and several sampling strategies across four standard image benchmark datasets (MNIST, SVHN, CIFAR-10, and STL-10) across model zoos containing 1,000 models each.
The analysis yields several key findings. First, layer-wise loss normalization resolves severe reconstruction failures where standard autoencoders collapsed into random guessing (about 10% accuracy on SVHN); with normalization, decoded models achieve functional baseline performance (approximately 51.5% initial accuracy). Second, models generated using density estimation over the top 30% performing weights learn substantially faster during fine-tuning, reaching in 25 epochs higher accuracy than scratch-trained models achieve in 50 epochs (e.g., about 74.2% versus 70.7% on SVHN). Third, sampled models exhibit sufficient diversity to build high-performing model ensembles at negligible extra computational cost, with 15-model ensembles reaching 77.6% accuracy on SVHN. Finally, the generated initializations generalize successfully to transfer learning across datasets and adapt effectively to unseen neural network architectures, such as networks with added residual skip connections.
These findings indicate that hyper-representations can act as versatile, data-free generative models for neural weights, significantly reducing the compute time, data requirements, and energy costs associated with training deep networks from scratch. Decision-makers and technical leaders can leverage these techniques to improve model initialization, accelerate multi-task learning, and aggregate knowledge across isolated internal model checkpoints without compromising data privacy.
Organizations should consider evaluating weight-space generation pipelines for rapid model prototyping and lightweight ensembling. However, stakeholders should note key limitations: the current evaluation relies on relatively small, uniform convolutional network architectures, and performance saturates on harder tasks (CIFAR-10 and STL-10) where low network capacity restricts total accuracy. Further research and pilot validations on modern, large-scale architectures are recommended before full deployment across production systems.
- Paper: HyperNetworks, David Ha et al. (2016). Introduces the foundational hypernetwork concept of using one neural network to generate and parameterize the weights of another network.
- Paper: Adversarial Autoencoders, Alireza Makhzani et al. (2015). Provides the foundational autoencoder-based generative modeling architecture and latent distribution regularization used to synthesize complex data points.
- Paper: Dreaming to Distill: Data-Free Knowledge Transfer via DeepInversion, Hongxu Yin et al. (2020). Establishes data-free knowledge transfer from pretrained neural network populations, motivating the source paper's data-free weight generation approach.
- Paper: Predicting Parameters in Deep Learning, Misha Denil et al. (2013). Demonstrates the underlying redundancy and predictability of neural network parameter spaces that makes learning weight-space generative models feasible.
- Paper: Tutorial on Variational Autoencoders, Carl Doersch (2016). Offers essential theoretical background on latent variable modeling and sampling techniques for continuous generative frameworks.
- Paper: Equivariant Architectures for Learning in Deep Weight Spaces, Aviv Navon et al. (2023). Extends deep weight-space learning by designing permutation-equivariant and invariant neural architectures explicitly tailored for processing neural network parameters.
- Paper: Adaptive Data-Free Quantization, Biao Qian et al. (2023). Applies data-free optimization to model compression and quantization by dynamically synthesizing calibration instances directly from internal network states.
- Paper: Up to 100x Faster Data-Free Knowledge Distillation, Gongfan Fang et al. (2022). Advances data-free model compression pipelines by utilizing meta-generators to synthesize diverse samples up to two orders of magnitude faster.
- Paper: Compressing Transformers: Features Are Low-Rank, but Weights Are Not!, Hao Yu et al. (2023). Investigates weight versus representation geometries in large networks, demonstrating how low-rank properties manifest across activations.
