Hyper-Representations as Generative Models: Sampling Unseen Neural Network Weights

Konstantin SchürholtBoris KnyazevXavier Giró-i-NietoDamian Borth

article2022NeurIPS78 citations

Introduces a generative approach using layer-wise loss normalization to sample diverse, high-performing neural network weights directly from model zoos for effective initialization, ensembling, and transfer learning without requiring underlying training data.

Listen

Organizations increasingly produce and store massive collections of trained machine learning models across online hubs, yet repurposing the underlying knowledge captured in these model populations remains a major challenge. Prior techniques either require direct access to the original proprietary training data or rely on representations that produce dysfunctional models when generating new network weights. The article addresses this gap by developing a generative framework that directly learns from populations of neural network weights—known as model zoos—without needing underlying data samples or class labels.

The main objective of the article is to demonstrate that an autoencoder framework can learn a compressed, regularized representation of neural network weights and reliably generate new, functional, and diverse neural network weights in a single computational pass. To evaluate this, the authors introduce a novel layer-wise loss normalization technique and several sampling strategies across four standard image benchmark datasets (MNIST, SVHN, CIFAR-10, and STL-10) across model zoos containing 1,000 models each.

The analysis yields several key findings. First, layer-wise loss normalization resolves severe reconstruction failures where standard autoencoders collapsed into random guessing (about 10% accuracy on SVHN); with normalization, decoded models achieve functional baseline performance (approximately 51.5% initial accuracy). Second, models generated using density estimation over the top 30% performing weights learn substantially faster during fine-tuning, reaching in 25 epochs higher accuracy than scratch-trained models achieve in 50 epochs (e.g., about 74.2% versus 70.7% on SVHN). Third, sampled models exhibit sufficient diversity to build high-performing model ensembles at negligible extra computational cost, with 15-model ensembles reaching 77.6% accuracy on SVHN. Finally, the generated initializations generalize successfully to transfer learning across datasets and adapt effectively to unseen neural network architectures, such as networks with added residual skip connections.

These findings indicate that hyper-representations can act as versatile, data-free generative models for neural weights, significantly reducing the compute time, data requirements, and energy costs associated with training deep networks from scratch. Decision-makers and technical leaders can leverage these techniques to improve model initialization, accelerate multi-task learning, and aggregate knowledge across isolated internal model checkpoints without compromising data privacy.

Organizations should consider evaluating weight-space generation pipelines for rapid model prototyping and lightweight ensembling. However, stakeholders should note key limitations: the current evaluation relies on relatively small, uniform convolutional network architectures, and performance saturates on harder tasks (CIFAR-10 and STL-10) where low network capacity restricts total accuracy. Further research and pilot validations on modern, large-scale architectures are recommended before full deployment across production systems.

  • Paper: HyperNetworks, David Ha et al. (2016). Introduces the foundational hypernetwork concept of using one neural network to generate and parameterize the weights of another network.
  • Paper: Adversarial Autoencoders, Alireza Makhzani et al. (2015). Provides the foundational autoencoder-based generative modeling architecture and latent distribution regularization used to synthesize complex data points.
  • Paper: Dreaming to Distill: Data-Free Knowledge Transfer via DeepInversion, Hongxu Yin et al. (2020). Establishes data-free knowledge transfer from pretrained neural network populations, motivating the source paper's data-free weight generation approach.
  • Paper: Predicting Parameters in Deep Learning, Misha Denil et al. (2013). Demonstrates the underlying redundancy and predictability of neural network parameter spaces that makes learning weight-space generative models feasible.
  • Paper: Tutorial on Variational Autoencoders, Carl Doersch (2016). Offers essential theoretical background on latent variable modeling and sampling techniques for continuous generative frameworks.
Cover for Hyper-Representations as Generative Models: Sampling Unseen Neural Network Weights

Abstract

Learning representations of neural network weights given a model zoo is an emerging and challenging area with many potential applications from model inspection, to neural architecture search or knowledge distillation. Recently, an autoencoder trained on a model zoo was able to learn a hyper-representation, which captures intrinsic and extrinsic properties of the models in the zoo. In this work, we extend hyper-representations for generative use to sample new model weights. We propose layer-wise loss normalization which we demonstrate is key to generate high-performing models and several sampling methods based on the topology of hyper-representations. The models generated using our methods are diverse, performant and capable to outperform strong baselines as evaluated on several downstream tasks: initialization, ensemble sampling and transfer learning. Our results indicate the potential of knowledge aggregation from model zoos to new models via hyper-representations thereby paving the avenue for novel research directions.

Table of Contents

  • 1 Introduction
  • 2 Background: Training Hyper-Representations
  • 3 Methods
  • 3.1 Layer-Wise Loss Normalization
  • 3.2 Sampling from Hyper-Representations
  • 3.2.1 Uniform S U
  • 3.2.2 Density estimation S KDE and counterfactual sampling S C
  • 3.2.3 Neighbor sampling S Neigh
  • 3.2.4 Latent space GAN S GAN
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Results
  • 4.2.1 Hyper-Representations are Robust and Smooth
  • 4.2.2 Sampling for In-dataset Initialization
  • 4.2.3 Sampling Initializations for Transfer Learning
  • 4.2.4 Sampling Initializations for Unseen Architectures
  • 4.3 Limitations of Zoos with Small Models
  • 5 Related Work
  • 6 Conclusion
  • Acknowledgments
  • References
  • Checklist

Knowls

  1. Knowl 1 — Layer-Wise Loss Normalization (LWLN) for Weight Autoencoders

    equation

    When training an autoencoder on neural network weight vectors w∈RN\mathbf{w} \in \mathbb{R}^N, standard Mean Squared Error (MSE) uniformly penalizes reconstruction errors across all weights and layers. Because standard network initialization schemes produce vastly different parameter scales and variances across layers, layers with smaller weight magnitudes collapse to the layer-wise mean under unnormalized MSE, rendering reconstructed networks non-functional (performance drops to random guessing). Layer-Wise Loss Normalization (LWLN) balances the reconstruction penalty across layers by weighting each layer's error by the inverse of its empirical variance over the training zoo:

    LˉMSE=1MN∑i=1M∑l=1L∥w^i(l)−wi(l)∥22σl2\bar{L}_{\text{MSE}} = \frac{1}{M N} \sum_{i=1}^{M} \sum_{l=1}^{L} \frac{\| \hat{\mathbf{w}}_i^{(l)} - \mathbf{w}_i^{(l)} \|_2^2}{\sigma_l^2}

    where MM is the number of neural network models in the dataset (model zoo), NN is the total parameter dimensionality of each network, LL is the number of layers in the architecture, wi(l)\mathbf{w}_i^{(l)} and w^i(l)\hat{\mathbf{w}}_i^{(l)} denote the ground-truth and reconstructed weight vectors for the ll-th layer of the ii-th model, and σl\sigma_l is the standard deviation of all weights in layer ll estimated across the training split of the model zoo.

  2. Knowl 2 — Self-Supervised Hyper-Representation Framework for Model Zoos

    model/method

    A hyper-representation autoencoder learns a low-dimensional manifold directly from a population of neural network parameter vectors {wi}i=1M⊂RN\{\mathbf{w}_i\}_{i=1}^M \subset \mathbb{R}^N without accessing training images or labels.

    Each model weight vector is structured as a sequence of token embeddings corresponding to individual convolutional or fully connected neurons. The encoder g:RN→RDg: \mathbb{R}^N \to \mathbb{R}^D processes this sequence through multi-head self-attention layers, summarizes the representation into a compression token, and applies a tanh⁡\tanh activation to yield a bounded latent embedding zi=g(wi)∈[−1,1]D\mathbf{z}_i = g(\mathbf{w}_i) \in [-1, 1]^D. The decoder h:RD→RNh: \mathbb{R}^D \to \mathbb{R}^N symmetrically decompresses zi\mathbf{z}_i, adds positional encodings, and reconstructs the weight vector w^i=h(zi)\hat{\mathbf{w}}_i = h(\mathbf{z}_i).

    The training loss is a multi-objective combination:

    L=βLˉMSE+(1−β)LcL = \beta \bar{L}_{\text{MSE}} + (1 - \beta) L_c

    where LˉMSE\bar{L}_{\text{MSE}} is the layer-wise loss-normalized reconstruction loss, LcL_c is a contrastive loss using weight permutations and random neuron erasing augmentations to enforce symmetry and compactness in latent space, and β∈[0,1]\beta \in [0, 1] balances the terms.

  3. Knowl 3 — Latent Space Weight Sampling via Dimension-Wise Kernel Density Estimation

    algorithm

    To sample novel functional neural network weights w∗=h(z∗)\mathbf{w}^* = h(\mathbf{z}^*) from a learned hyper-representation autoencoder without assuming a standard Gaussian prior, Kernel Density Estimation (SKDES_{\text{KDE}}) models the empirical distribution over the latent anchor embeddings {zi}i=1M⊂RD\{\mathbf{z}_i\}_{i=1}^M \subset \mathbb{R}^D.

    Under a conditional independence assumption across latent dimensions, p(z(j)∣z(∖j),w)=p(z(j)∣w)p(z^{(j)} \mid \mathbf{z}^{(\setminus j)}, \mathbf{w}) = p(z^{(j)} \mid \mathbf{w}), a separate 1D Gaussian KDE is fitted to each dimension j∈{1,…,D}j \in \{1, \dots, D\}:

    p(z(j))=1Mhbw∑i=1MK(z(j)−zi(j)hbw),K(u)=12πexp⁡(−u22)p(z^{(j)}) = \frac{1}{M h_{\text{bw}}} \sum_{i=1}^{M} K\left(\frac{z^{(j)} - z_i^{(j)}}{h_{\text{bw}}}\right), \quad K(u) = \frac{1}{\sqrt{2\pi}} \exp\left(-\frac{u^2}{2}\right)

    where hbwh_{\text{bw}} is the bandwidth hyperparameter.

    Input: Trained decoder hh, training embeddings {zi}i=1M\{\mathbf{z}_i\}_{i=1}^M, bandwidth hbwh_{\text{bw}}
    Output: Generated network weight vector w∗\mathbf{w}^*
    for each latent dimension j=1j = 1 to DD do
        Fit 1D Gaussian KDE p(z(j))p(z^{(j)}) on anchor values {zi(j)}i=1M\{z_i^{(j)}\}_{i=1}^M using bandwidth hbwh_{\text{bw}}
        Sample z∗(j)∼p(z(j))z^{*(j)} \sim p(z^{(j)})
    end for
    Construct latent vector z∗=[z∗(1),z∗(2),…,z∗(D)]\mathbf{z}^* = [z^{*(1)}, z^{*(2)}, \dots, z^{*(D)}]
    Decode weight vector w∗=h(z∗)\mathbf{w}^* = h(\mathbf{z}^*)
    return w∗\mathbf{w}^*

    In the quality-focused variant SKDE30S_{\text{KDE30}}, the anchor set is filtered prior to KDE fitting to contain only the embeddings of the top 30% best-performing models in the training zoo.

  4. Knowl 4 — Latent Space Sampling via UMAP Neighborhood Inversion and Latent GAN

    model/method

    Due to high sparsity in the DD-dimensional hyper-representation space, two alternative sampling strategies model latent topology:

    1. Neighborhood-Based Inversion (SNeighS_{\text{Neigh}} / SNeigh30S_{\text{Neigh30}}): A dimensionality reduction mapping k:RD→Rdk: \mathbb{R}^D \to \mathbb{R}^d (with d=3d = 3) based on UMAP projects anchor embeddings zi\mathbf{z}_i into low-dimensional points ni=k(zi)\mathbf{n}_i = k(\mathbf{z}_i). A point n∗\mathbf{n}^* is sampled uniformly from the bounding hypercube U(min⁡(n),max⁡(n))\mathcal{U}(\min(\mathbf{n}), \max(\mathbf{n})), and mapped back to the hyper-representation space via approximate inverse UMAP: z∗=k−1(n∗)\mathbf{z}^* = k^{-1}(\mathbf{n}^*). Weights are decoded via w∗=h(z∗)\mathbf{w}^* = h(\mathbf{z}^*).

    2. Latent Space GAN (SGANS_{\text{GAN}} / SGAN30S_{\text{GAN30}}): A generative adversarial network generator G:Rdnoise→RDG: \mathbb{R}^{d_{\text{noise}}} \to \mathbb{R}^D (with dnoise=16d_{\text{noise}} = 16) is trained directly on the latent anchor representations {zi}i=1M\{\mathbf{z}_i\}_{i=1}^M. Latent samples are generated from Gaussian noise n∗∼N(0,I)\mathbf{n}^* \sim \mathcal{N}(0, I) via z∗=G(n∗)\mathbf{z}^* = G(\mathbf{n}^*) and decoded into network parameters w∗=h(z∗)\mathbf{w}^* = h(\mathbf{z}^*).

    In both methods, the top 30% variants (SNeigh30S_{\text{Neigh30}} and SGAN30S_{\text{GAN30}}) train exclusively on embeddings of the top 30% performing models in the zoo.

  5. Knowl 5 — In-Dataset Initialization and Fine-Tuning Performance Across Datasets

    data/table

    Models generated by decoding latent samples drawn with SKDE30S_{\text{KDE30}} from an autoencoder trained with layer-wise loss normalization (LWLN) yield functional initializations at epoch 0 that far exceed random guessing (~10%) across datasets, whereas the baseline autoencoder without LWLN (BKDE30B_{\text{KDE30}}) collapses to random guessing on SVHN, CIFAR-10, and STL-10. When fine-tuned, generated models learn faster than networks trained from scratch with SGD (BTB_T), achieving higher accuracy in 25 epochs than BTB_T achieves in 50 epochs on MNIST and SVHN.

    Method Epoch MNIST SVHN CIFAR-10 STL-10
    BTB_T 0 ≈10%\approx 10\% ≈10%\approx 10\% ≈10%\approx 10\% ≈10%\approx 10\%
    BKDE30B_{\text{KDE30}} 0 63.2±7.263.2 \pm 7.2 10.1±3.210.1 \pm 3.2 15.5±3.415.5 \pm 3.4 12.7±3.412.7 \pm 3.4
    SKDE30S_{\text{KDE30}} 0 68.6±6.7\mathbf{68.6 \pm 6.7} 51.5±5.9\mathbf{51.5 \pm 5.9} 26.9±4.9\mathbf{26.9 \pm 4.9} 19.7±2.1\mathbf{19.7 \pm 2.1}
    BTB_T 1 20.6±1.620.6 \pm 1.6 19.4±0.619.4 \pm 0.6 27.5±2.127.5 \pm 2.1 15.4±1.815.4 \pm 1.8
    BKDE30B_{\text{KDE30}} 1 83.2±1.283.2 \pm 1.2 67.4±2.067.4 \pm 2.0 39.7±0.639.7 \pm 0.6 26.4±1.6\mathbf{26.4 \pm 1.6}
    SKDE30S_{\text{KDE30}} 1 83.7±1.3\mathbf{83.7 \pm 1.3} 69.9±1.6\mathbf{69.9 \pm 1.6} 44.0±0.5\mathbf{44.0 \pm 0.5} 25.9±1.625.9 \pm 1.6
    BTB_T 25 83.3±2.683.3 \pm 2.6 66.7±8.566.7 \pm 8.5 46.1±1.346.1 \pm 1.3 35.0±1.335.0 \pm 1.3
    BKDE30B_{\text{KDE30}} 25 93.2±0.6\mathbf{93.2 \pm 0.6} 75.4±0.9\mathbf{75.4 \pm 0.9} 48.1±0.648.1 \pm 0.6 38.4±0.9\mathbf{38.4 \pm 0.9}
    SKDE30S_{\text{KDE30}} 25 93.0±0.793.0 \pm 0.7 74.2±1.474.2 \pm 1.4 48.6±0.5\mathbf{48.6 \pm 0.5} 38.1±1.138.1 \pm 1.1
    BTB_T 50 91.1±2.691.1 \pm 2.6 70.7±8.870.7 \pm 8.8 48.7±1.448.7 \pm 1.4 39.0±1.039.0 \pm 1.0

    Values represent mean ±\pm standard deviation of test classification accuracy (%) evaluated over populations of at least 50 models.

  6. Knowl 6 — Rapid Ensemble Generation and Weight Optimization Divergence

    empirical result

    Sampling diverse model weight vectors from hyper-representations enables the generation of high-performing model ensembles at negligible computational cost:

    1. Ensemble Performance: On SVHN, an ensemble composed of 15 models generated using SKDE30S_{\text{KDE30}} reaches an average test accuracy of 77.6%. This outperforms single models trained with SGD for 50 epochs (70.7%) and baseline autoencoder ensembles (BKDE30B_{\text{KDE30}} without LWLN), which stagnate below 20% accuracy regardless of ensemble size.

    2. Optimization Divergence: Reconstructed weights w^=h(g(w))\hat{\mathbf{w}} = h(g(\mathbf{w})) initialized from models trained for 25 epochs do not merely mirror their originals (w)(\mathbf{w}) when fine-tuned. Instead, w^\hat{\mathbf{w}} moves further apart from w\mathbf{w} in Euclidean distance ∥w−w^∥22/∥w∥22\|\mathbf{w} - \hat{\mathbf{w}}\|_2^2 / \|\mathbf{w}\|_2^2 and cosine distance 1−cos⁡(w,w^)1 - \cos(\mathbf{w}, \hat{\mathbf{w}}), exploring a different trajectory on the loss surface that converges faster and achieves higher test accuracy.

  7. Knowl 7 — Transfer Learning Weight Initialization Across Dataset Domains

    data/table

    Hyper-representations trained on a source image domain can generate effective initialization weights for training on a different target domain, outperforming models trained from scratch (BTB_T) and competing favorably with standard pre-trained fine-tuning (BFB_F).

    SVHN →\to MNIST STL-10 →\to CIFAR-10
    Method Ep. 0 Ep. 1 Ep. 50 Ep. 0 Ep. 1 Ep. 50
    BTB_T 10.0±0.610.0 \pm 0.6 20.6±1.620.6 \pm 1.6 91.1±1.091.1 \pm 1.0 10.1±1.310.1 \pm 1.3 27.5±2.127.5 \pm 2.1 48.7±1.448.7 \pm 1.4
    BFB_F 33.4±5.433.4 \pm 5.4 84.4±7.484.4 \pm 7.4 95.0±0.895.0 \pm 0.8 15.3±2.315.3 \pm 2.3 29.4±1.929.4 \pm 1.9 49.2±0.7\mathbf{49.2 \pm 0.7}
    SKDE30S_{\text{KDE30}} 31.8±5.6\mathbf{31.8 \pm 5.6} 86.9±1.4\mathbf{86.9 \pm 1.4} 95.5±0.4\mathbf{95.5 \pm 0.4} 14.5±1.914.5 \pm 1.9 29.6±2.0\mathbf{29.6 \pm 2.0} 48.8±0.948.8 \pm 0.9
    SNeigh30S_{\text{Neigh30}} 10.7±2.710.7 \pm 2.7 79.2±3.379.2 \pm 3.3 95.5±0.7\mathbf{95.5 \pm 0.7} 10.1±2.110.1 \pm 2.1 29.2±1.929.2 \pm 1.9 48.9±0.748.9 \pm 0.7
    SGAN30S_{\text{GAN30}} 10.4±2.410.4 \pm 2.4 75.0±6.375.0 \pm 6.3 94.9±0.794.9 \pm 0.7 10.2±2.510.2 \pm 2.5 28.6±1.828.6 \pm 1.8 48.8±0.848.8 \pm 0.8

    Values report mean ±\pm standard deviation of test classification accuracy (%) on the target dataset. For SVHN →\to MNIST, generated initializations from SKDE30S_{\text{KDE30}} enable rapid early learning (86.9%86.9\% at epoch 1 vs 20.6%20.6\% for BTB_T) and higher final accuracy at epoch 50 (95.5%95.5\% vs 91.1%91.1\% for BTB_T and 95.0%95.0\% for BFB_F).

  8. Knowl 8 — Zero-Shot Weight Conditioning on Unseen Model Zoos

    data/table

    A hyper-representation autoencoder trained on the weight zoo of one dataset can encode and reconstruct weight vectors from an entirely unseen dataset zoo without any model retraining, generating weights that surpass random guessing (10%).

    Training Zoo Conditioning Zoo (Unseen) One Model (Mean / Max) Ensemble (Mean / Max)
    MNIST SVHN 12.7%/19.8%12.7 \% / 19.8 \% 13.4%/18.7%13.4 \% / 18.7 \%
    SVHN MNIST 16.2%/26.0%16.2 \% / 26.0 \% 22.1%/29.8%22.1 \% / 29.8 \%
    CIFAR-10 STL-10 18.0%/24.4%18.0 \% / 24.4 \% 23.8%/26.7%23.8 \% / 26.7 \%
    STL-10 CIFAR-10 16.3%/21.2%16.3 \% / 21.2 \% 20.0%/23.0%20.0 \% / 23.0 \%

    Ensembling multiple reconstructed models consistently boosts accuracy on the unseen domain (e.g., reaching up to 29.8%29.8\% on MNIST when conditioned through an autoencoder trained strictly on SVHN weights).

  9. Knowl 9 — Cross-Architecture Weight Transfer and Initialization

    data/table

    Weights generated from a hyper-representation trained on a standard 3-layer convolutional architecture (3-conv + 2-FC) on MNIST can initialize structurally different, unseen architectures trained on SVHN. Parameter mismatches are resolved either by re-distributing weights across layers or by randomly initializing additional parameters (e.g., 1×11\times 1 residual projections):

    Architecture Initialization Epoch 1 Epoch 5 Epoch 50
    3-conv (r. i.) + res-skip (r. i.) 18.9±1.618.9 \pm 1.6 31.4±1731.4 \pm 17 50.6±2850.6 \pm 28
    3-conv (gen.) + res-skip (r. i.) 34.5±14\mathbf{34.5 \pm 14} 60.5±21\mathbf{60.5 \pm 21} 68.0±21\mathbf{68.0 \pm 21}
    4-conv (r. i.) 19.2±1.019.2 \pm 1.0 19.2±0.919.2 \pm 0.9 55.2±1155.2 \pm 11
    4-conv (gen.) 44.0±4.5\mathbf{44.0 \pm 4.5} 57.8±3.5\mathbf{57.8 \pm 3.5} 67.6±1.9\mathbf{67.6 \pm 1.9}
    4-conv + id.-skip (r. i.) 18.9±1.018.9 \pm 1.0 19.6±1.719.6 \pm 1.7 56.4±7.956.4 \pm 7.9
    4-conv + id.-skip (gen.) 48.0±4.0\mathbf{48.0 \pm 4.0} 59.9±2.5\mathbf{59.9 \pm 2.5} 66.4±1.7\mathbf{66.4 \pm 1.7}

    Here r. i. denotes random initialization and gen. denotes generated weights sampled via SKDE30S_{\text{KDE30}}. Generated weight initializations reach higher accuracy in 5 epochs than randomly initialized architectures reach after 50 epochs across all structural variations.

  10. Knowl 10 — Capacity Saturation in Small-Scale Model Zoos

    limitation

    When model zoos are built using low-capacity architectures (e.g., 3 convolutional and 2 fully-connected layers) on complex image benchmarks such as CIFAR-10 and STL-10, the accuracy of all models in the zoo saturates early (at ≈50%\approx 50\% for CIFAR-10 and ≈40%\approx 40\% for STL-10). Because the base models have high residual loss throughout training, the stochastic gradient descent weight updates remain large and noisy without achieving true convergence. Consequently, the learned weight distributions contain high noise and a low signal-to-noise ratio, preventing generated and fine-tuned models from significantly surpassing the performance ceiling of standard baselines.

Coverage note — None was omitted; all primary methodological contributions (LWLN, KDE/UMAP/GAN sampling strategies), core theoretical motivations, downstream experimental evaluations (in-dataset, transfer learning, zero-shot conditioning, cross-architecture transfer), and stated limitations are fully represented.

References

  1. 1.Luca Bertinetto, João F Henriques, Jack Valmadre, Philip Torr, and Andrea Vedaldi. Learning feed-forward one-shot learners. Advances in Neural Information Processing Systems, 29, 2016. 10
  2. 2.Arantxa Casanova, Marlène Careil, Jakob Verbeek, Michal Drozdzal, and Adriana Romero Soriano. Instance-conditioned gan. Advances in Neural Information Processing Systems, 34, 2021. 8
  3. 3.Tianqi Chen, Ian Goodfellow, and Jonathon Shlens. Net2Net: Accelerating Learning via Knowledge Transfer. In International Conference on Learning Representations (ICLR), April 2016. doi: 10.48550/arXiv.1511.05641. 9
  4. 4.Wei-Yu Chen, Yen-Cheng Liu, Zsolt Kira, Yu-Chiang Frank Wang, and Jia-Bin Huang. A closer look at few-shot classification. International Conference on Learning Representations (ICLR), 2019. 5, 10
  5. 5.Adam Coates, Honglak Lee, and Andrew Y Ng. An Analysis of Single-Layer Networks in Unsupervised Feature Learning. In Proceedings of the 14th International Con- Ference on Artificial Intelligence and Statistics (AISTATS), page 9, 2011. 5
  6. 6.Yann N Dauphin and Samuel Schoenholz. MetaInit: Initializing learning by learning to initialize. In Neural Information Processing Systems, page 13, 2019. 10
  7. 7.Misha Denil, Babak Shakibi, Laurent Dinh, and Marc’Aurelio Ranzato. Predicting Parameters in Deep Learning. In Neural Information Processing Systems (NeurIPS), page 9, 2013. 10
  8. 8.Lior Deutsch. Generating Neural Networks with Neural Networks. arXiv:1801.01952 [cs, stat], April 2018. 2, 10
  9. 9.Guneet S Dhillon, Pratik Chaudhari, Avinash Ravichandran, and Stefano Soatto. A baseline for few-shot image classification. International Conference on Learning Representations (ICLR), 2019. 5, 10
  10. 10.Partha Ghosh, Mehdi S. M. Sajjadi, Antonio Vergari, Michael Black, and Bernhard Schölkopf. From Variational to Deterministic Autoencoders. In arXiv:1903.12436 [Cs, Stat], May 2020. 3, 4, 21
  11. 11.Xavier Glorot and Yoshua Bengio. Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, PMLR, page 8, 2010. 3, 5
  12. 12.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative Adversarial Nets. In Conference on Neural Information Processing Systems (NeurIPS), page 9, 2014. 4
  13. 13.Yong Guo, Qi Chen, Jian Chen, Qingyao Wu, Qinfeng Shi, and Mingkui Tan. Auto-embedding generative adversarial networks for high resolution image synthesis. IEEE Transactions on Multimedia, 21(11):2726–2737, 2019. 3, 4
  14. 14.David Ha, Andrew Dai, and Quoc V. Le. HyperNetworks. In arXiv:1609.09106 [Cs], 2016. 2, 10
  15. 15.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification. In arXiv:1502.01852 [Cs], 2015. 3, 5
  16. 16.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep Residual Learning for Image Recognition. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016. 9
  17. 17.Kaiming He, Ross Girshick, and Piotr Dollár. Rethinking imagenet pre-training. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4918–4927, 2019. 6
  18. 18.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems, 30, 2017. 21
  19. 19.Li Jing, Pascal Vincent, Yann LeCun, and Yuandong Tian. Understanding Dimensional Collapse in Contrastive Self-supervised Learning. In International Conference on Learning Representations, September 2021. 19
  20. 20.Diederik P. Kingma and Max Welling. Auto-Encoding Variational Bayes. In International Conference on Learning Representations (ICLR), 2013. 3, 21
  21. 21.Boris Knyazev, Michal Drozdzal, Graham W. Taylor, and Adriana Romero-Soriano. Parameter Prediction for Unseen Deep Architectures. In Conference on Neural Information Processing Systems (NeurIPS), 2021. 2, 3, 10
  22. 22.Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby. Big transfer (bit): General visual representation learning. In European conference on computer vision, pages 491–507. Springer, 2020. 5, 10
  23. 23.Alex Krizhevsky. Learning Multiple Layers of Features from Tiny Images. page 60, 2009. 2, 5
  24. 24.Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner. Gradient-Based Learning Applied to Document Recognition. Proceedings of the IEEE, 86(11):2278–2324, November 1998. 5
  25. 25.Iou-Jen Liu, Jian Peng, and Alexander G. Schwing. Knowledge Flow: Improve Upon Your Teachers. In International Conference on Learning Representations (ICLR), April 2019. 10
  26. 26.Jinlin Liu, Yuan Yao, and Jianqiang Ren. An acceleration framework for high resolution image synthesis. arXiv preprint arXiv:1909.03611, 2019. 3, 4
  27. 27.Charles H Martin, Tongsu Serena Peng, and Michael W Mahoney. Predicting trends in the quality of state-of-the-art neural networks without access to training or testing data. Nature Communications, 12(1):1–13, 2021. 1
  28. 28.Leland McInnes, John Healy, and Nathaniel Saul. UMAP: Uniform Manifold Approximation and Projection. 2018. 4
  29. 29.Thomas Mensink, Jasper Uijlings, Alina Kuznetsova, Michael Gygli, and Vittorio Ferrari. Factors of Influence for Transfer Learning across Diverse Appearance Domains and Task Types. arXiv:2103.13318 [cs], November 2021. 8, 10
  30. 30.Dmytro Mishkin and Jiri Matas. All you need is a good init. In International Conference on Learning Representations (ICLR). arXiv, 2016. doi: 10.48550/arXiv.1511.06422. 6
  31. 31.Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida. Spectral Normalization for Generative Adversarial Networks. In International Conference on Learning Representations (ICLR), 2018. 21
  32. 32.Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng. Reading Digits in Natural Images with Unsupervised Feature Learning. In NIPS Workshop on Deep Learning and Unsupervised Feature Learning 2011, page 9, 2011. 3, 5
  33. 33.William Peebles, Ilija Radosavovic, Tim Brooks, Alexei A. Efros, and Jitendra Malik. Learning to Learn with Generative Models of Neural Network Checkpoints, September 2022. 10
  34. 34.Antti Rasmus, Mathias Berglund, Mikko Honkala, Harri Valpola, and Tapani Raiko. Semisupervised learning with ladder networks. Advances in neural information processing systems, 28, 2015. 6
  35. 35.Neale Ratzlaff and Li Fuxin. HyperGAN: A Generative Model for Diverse, Performant Neural Networks. In Proceedings of the 36th International Conference on Machine Learning, pages 5361–5369. PMLR, May 2019. 2, 4, 10
  36. 36.James Requeima, Jonathan Gordon, John Bronskill, Sebastian Nowozin, and Richard E Turner. Fast and flexible multi-task classification using conditional neural adaptive processes. Advances in Neural Information Processing Systems, 32, 2019. 10
  37. 37.Konstantin Schürholt, Dimche Kostadinov, and Damian Borth. Self-Supervised Representation Learning on Neural Network Weights for Model Characteristic Prediction. In Conference on Neural Information Processing Systems (NeurIPS), volume 35, 2021. 1, 2, 3, 4, 5, 9, 10, 16, 21
  38. 38.Konstantin Schürholt, Diyar Taskiran, Boris Knyazev, Xavier Giró-i-Nieto, and Damian Borth. Model Zoos: A Dataset of Diverse Populations of Neural Network Models. In Thirty-Sixth Conference on Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track, September 2022. 5, 9, 16
  39. 39.Yang Shu, Zhi Kou, Zhangjie Cao, Jianmin Wang, and Mingsheng Long. Zoo-Tuning: Adaptive Transfer from a Zoo of Models. In International Conference on Machine Learning (ICML), page 12, 2021. 10
  40. 40.Thomas Unterthiner, Daniel Keysers, Sylvain Gelly, Olivier Bousquet, and Ilya Tolstikhin. Predicting Neural Network Accuracy from Weights. arXiv:2002.11448 [cs, stat], February 2020. 1
  41. 41.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention Is All You Need. In arXiv:1706.03762 [Cs], December 2017. 2
  42. 42.Jiayun Wang, Yubei Chen, Stella X. Yu, Brian Cheung, and Yann LeCun. Recurrent Parameter Generators, July 2021. 9
  43. 43.Kuan-Chieh Wang, Paul Vicol, James Lucas, Li Gu, Roger Grosse, and Richard Zemel. Adversarial distillation of bayesian neural network posteriors. In International conference on machine learning, pages 5190–5199. PMLR, 2018. 10
  44. 44.Jason Yosinski, Jeff Clune, Yoshua Bengio, and Hod Lipson. How transferable are features in deep neural networks? In Neural Information Processing Systems (NeurIPS), November 2014. 10
  45. 45.Chris Zhang, Mengye Ren, and Raquel Urtasun. Graph HyperNetworks for Neural Architecture Search. In International Conference on Learning Representations (ICLR), 2019. 2, 10
  46. 46.Andrey Zhmoginov, Mark Sandler, and Max Vladymyrov. HyperTransformer: Model Generation for Supervised and Semi-Supervised Few-Shot Learning. In International Conference on Machine Learning (ICML), January 2022. 2, 10
  47. 47.Chen Zhu, Renkun Ni, Zheng Xu, Kezhi Kong, W Ronny Huang, and Tom Goldstein. Gradinit: Learning to initialize neural networks for stable and efficient training. Advances in Neural Information Processing Systems, 34, 2021. 10
  48. 48.Fuzhen Zhuang, Zhiyuan Qi, Keyu Duan, Dongbo Xi, Yongchun Zhu, Hengshu Zhu, Hui Xiong, and Qing He. A Comprehensive Survey on Transfer Learning. In Proceedings of IEEE, 2020. 10

Citation

MLA
Schürholt, K., et al. “Hyper-Representations as Generative Models: Sampling Unseen Neural Network Weights”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 27906–20, https://proceedings.neurips.cc/paper_files/paper/2022/file/b2c4b7d34b3d96b9dc12f7bce424b7ae-Paper-Conference.pdf.
APA
Schürholt, K., Knyazev, B., Giró-i-Nieto, X., & Borth, D. (2022). Hyper-Representations as Generative Models: Sampling Unseen Neural Network Weights. Advances in Neural Information Processing Systems, 35, 27906–27920. https://proceedings.neurips.cc/paper_files/paper/2022/file/b2c4b7d34b3d96b9dc12f7bce424b7ae-Paper-Conference.pdf
Chicago
Schürholt, K., B. Knyazev, X. Giró-i-Nieto, and D. Borth. 2022. “Hyper-Representations as Generative Models: Sampling Unseen Neural Network Weights”. Advances in Neural Information Processing Systems 35: 27906–20. https://proceedings.neurips.cc/paper_files/paper/2022/file/b2c4b7d34b3d96b9dc12f7bce424b7ae-Paper-Conference.pdf.
Harvard
Schürholt, K. et al. (2022) “Hyper-Representations as Generative Models: Sampling Unseen Neural Network Weights”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 27906–27920. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/b2c4b7d34b3d96b9dc12f7bce424b7ae-Paper-Conference.pdf.
Vancouver
1. Schürholt K, Knyazev B, Giró-i-Nieto X, Borth D (2022) Hyper-Representations as Generative Models: Sampling Unseen Neural Network Weights. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 27906–27920

BibTeX

@inproceedings{schurholt2022hyper,
  title = {Hyper-Representations as Generative Models: Sampling Unseen Neural Network Weights},
  author = {Schürholt, Konstantin and Knyazev, Boris and Giró-i-Nieto, Xavier and Borth, Damian},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {27906-27920},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/b2c4b7d34b3d96b9dc12f7bce424b7ae-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors