Built independently by an author, for readers. Read the story and support ChapterPal

keyword

hyper-representation learning

Hyper-representation learning is a machine learning paradigm focused on learning low-dimensional, structured representations of the parameter spaces, such as weights and biases, across collections or populations of trained neural networks. Instead of extracting features from standard data modalities like images or text, this approach treats entire neural network models as data samples to uncover the underlying manifold and organization of their weight space. These learned representations capture both intrinsic attributes, such as architecture configurations and training hyperparameters, and extrinsic characteristics, including task performance and generalization capability. By modeling the relationships between network parameters and their functional behaviors, hyper-representation learning enables tasks such as model property prediction, neural network inspection, transfer learning, and the direct generative synthesis or initialization of new neural network weights.

1 item

Hyper-Representations as Generative Models: Sampling Unseen Neural Network Weights

Hyper-Representations as Generative Models: Sampling Unseen Neural Network Weights

Konstantin Schürholt, Boris Knyazev, Xavier Giró-i-Nieto, Damian Borth

OrganizationsAIML Lab, School of Computer ScienceInstitut de Robòtica i Informàtica Industrial, CSIC-UPCSamsung SAILUniversitat Politècnica de CatalunyaUniversity of St. Gallen

Why you should read this

Introduces a generative approach using layer-wise loss normalization to sample diverse, high-performing neural network weights directly from model zoos for effective initialization, ensembling, and transfer learning without requiring underlying training data.

Learning representations of neural network weights given a model zoo is an emerging and challenging area with many potential applications from model inspection, to neural architecture search or knowledge distillation. Recently, an autoencoder trained on a model zoo was able to learn a hyper-representation, which captures intrinsic and extrinsic properties of the models in the zoo. In this work, we extend hyper-representations for generative use to sample new model weights. We propose layer-wise loss normalization which we demonstrate is key to generate high-performing models and several sampling methods based on the topology of hyper-representations. The models generated using our methods are diverse, performant and capable to outperform strong baselines as evaluated on several downstream tasks: initialization, ensemble sampling and transfer learning. Our results indicate the potential of knowledge aggregation from model zoos to new models via hyper-representations thereby paving the avenue for novel research directions.

Added

2026-09-26