Semi-supervised Learning with Ladder Networks

Antti RasmusMathias BerglundM. HonkalaHarri ValpolaT. Raiko

article2015NeurIPS1,433 citations

Introduces a semi-supervised Ladder Network architecture that trains supervised and unsupervised reconstruction targets simultaneously via standard backpropagation, drastically reducing the labeled data needed for high-accuracy image classification without layer-wise pre-training.

Listen

Modern machine learning systems typically require massive volumes of manually labeled data to achieve high performance. In many practical industries, acquiring these labels is prohibitively expensive, time-consuming, or requires scarce domain expertise, while raw, unlabeled data remains abundant. Traditional semi-supervised methods attempt to use unlabeled data to assist training, but they often struggle to scale to deep architectures or require complex, multi-stage training pipelines that separate feature learning from classification.

The article demonstrates that combining supervised classification with an auxiliary, layer-wise unsupervised denoising task inside a single neural network architecture—termed a Ladder network—significantly improves classification accuracy when labeled data is scarce. It evaluates this framework across standard benchmark image datasets under varying levels of label availability, demonstrating how a feedforward neural network can simultaneously learn invariant abstract representations and class boundaries in a unified training pass.

To evaluate this capability, the authors constructed deep fully connected and convolutional architectures augmented with an auxiliary decoding pathway. The network processes both clean and intentionally corrupted versions of an input, using lateral skip connections to reconstruct representations at every layer while simultaneously predicting class labels. Experiments were conducted on the MNIST digit classification benchmark (under permutation-invariant and convolutional setups with as few as 100 labeled samples) and the CIFAR-10 image dataset (using 4,000 labeled samples). Optimization was performed in an end-to-end manner using standard backpropagation and batch normalization without requiring layer-by-layer pre-training.

The findings show substantial performance improvements across all tested configurations. On the permutation-invariant MNIST task with only 100 labeled examples, the full Ladder network achieved an error rate of 1.06%, drastically outperforming prior semi-supervised methods which had error rates ranging between 2.12% and 16.86%, and improving on the purely supervised baseline error of 21.74%. When evaluated with full labels on the same benchmark, the method established a new record error rate of 0.57%. On the more complex CIFAR-10 image dataset with 4,000 labels, a simplified top-layer variant of the architecture reduced classification error from a heavily regularized supervised baseline of 23.33% down to 20.40%.

These results indicate that organizations can achieve high-accuracy predictive models with substantially smaller investments in manual data labeling, reducing project costs and shortening deployment timelines. Because the unsupervised decoder acts as an effective regularizer, models can scale to larger parameter capacities without overfitting. Furthermore, the approach integrates into standard deep learning workflows with minimal overhead, approximately tripling computation per training pass but often converging faster due to more efficient data utilization.

Organizations developing computer vision and classification systems should consider adopting Ladder network architectures, particularly when entering new domains where labeled data is scarce. For rapid deployment, engineering teams can implement the simplified top-level denoising variant (the Gamma-model) into existing feedforward pipelines without designing a complete multi-layer decoder. Future technical initiatives should focus on exploring asymmetric encoder-decoder configurations and extending the architecture to sequential domains, such as video analysis, where manual annotation is especially costly.

Decision-makers should note certain operational boundaries and variances. The simplified top-level variant exhibited occasional convergence instability and confirmation bias when trained on extremely sparse label sets, where approximately 5% of runs produced outlier error rates above 2%. Confidence in the findings is high for standard image classification benchmarks, but teams should conduct pilot validations when porting the architecture to noisy real-world data or novel data modalities.

arXiv: 1507.02672
Cover for Semi-supervised Learning with Ladder Networks

Abstract

We combine supervised learning with unsupervised learning in deep neural networks. The proposed model is trained to simultaneously minimize the sum of supervised and unsupervised cost functions by backpropagation, avoiding the need for layer-wise pre-training. Our work builds on the Ladder network proposed by Valpola (2015), which we extend by combining the model with supervision. We show that the resulting model reaches state-of-the-art performance in semi-supervised MNIST and CIFAR-10 classification, in addition to permutation-invariant MNIST classification with all labels.

Table of Contents

  • 1 Introduction
  • 2 Derivation and justification
  • 3 Implementation of the Model
  • 3.1 General Steps for Implementing the Ladder Network
  • 3.2 Fully Connected MLP as Encoder
  • 3.3 Decoder for Unsupervised Learning
  • 3.4 Variations
  • 4 Experiments
  • 4.1 MNIST dataset
  • 4.1.1 Fully connected MLP
  • 4.1.2 Convolutional networks
  • 4.2 Convolutional networks on CIFAR-10
  • 5 Related Work
  • 6 Discussion
  • References
  • A Specification of the convolutional models
  • B Formulation of the Denoising Function

Knowls

  1. Knowl 1 — Ladder Network Architecture for Semi-Supervised Learning

    model/method

    The Ladder network is a neural network architecture designed for semi-supervised learning that combines a supervised feedforward classifier with an auxiliary layer-wise unsupervised denoising autoencoder. The system consists of three main computational pathways:

    1. Clean Encoder: A standard feedforward pathway that takes an uncorrupted input xx and computes layer activations z(l)z^{(l)} and h(l)h^{(l)} up to class predictions y=h(L)y = h^{(L)}. The intermediate representations z(l)z^{(l)} serve as clean reconstruction targets for the decoder.

    2. Corrupted Encoder: A parallel pathway that processes a noisy input x~=x+n(0)\tilde{x} = x + n^{(0)} (with additive isotropic Gaussian noise) and injects noise n(l)n^{(l)} at every hidden preactivation z~(l)\tilde{z}^{(l)} after batch normalization, yielding corrupted activations h~(l)\tilde{h}^{(l)} and output y~\tilde{y}. Supervised classification loss is computed on y~\tilde{y} for labeled training examples.

    3. Decoder with Lateral Connections: A top-down pathway that inverts the encoder mappings. At each layer ll, the decoder receives a top-down prior u(l)u^{(l)} from the layer above z^(l+1)\hat{z}^{(l+1)} via a batch-normalized projection u(l)=NB(V(l+1)z^(l+1))u^{(l)} = \text{NB}(V^{(l+1)} \hat{z}^{(l+1)}), and combines it laterally with the corrupted lateral activation z~(l)\tilde{z}^{(l)} through a local denoising function z^(l)=g(z~(l),u(l))\hat{z}^{(l)} = g(\tilde{z}^{(l)}, u^{(l)}).

    The skip connections between the corrupted encoder and the decoder allow lower levels to handle low-level detail reconstruction, freeing higher layers to represent abstract, invariant features relevant to classification. The model is trained end-to-end using standard backpropagation on the combined supervised and layer-wise unsupervised denoising objectives across both labeled and unlabeled data.

  2. Knowl 2 — Ladder Network Training and Evaluation Procedure

    algorithm

    The Ladder network processes minibatches of input samples x(n)x(n) (labeled with target class t(n)t(n) or unlabeled) by executing a corrupted forward pass, a clean forward pass to generate denoising targets, a top-down decoder reconstruction pass, and computing the sum of supervised cross-entropy and layer-wise mean squared reconstruction errors.

    Input: Training sample x(n)x(n), optional target label t(n)t(n), network depth LL, weights W(l)W^{(l)} and V(l)V^{(l)}, batch normalization scale and bias parameters γ(l)\gamma^{(l)} and β(l)\beta^{(l)}, cost weights λl\lambda_l, noise distributions N(0,σ2)\mathcal{N}(0, \sigma^2)
    Output: Classification distribution P(y∣x)P(y \mid x), total loss CC
    # 1. Corrupted Encoder Pass
    h~(0)←z~(0)←x(n)+n(0)\tilde{h}^{(0)} \leftarrow \tilde{z}^{(0)} \leftarrow x(n) + n^{(0)}
    for l=1l = 1 to LL do
        z~pre(l)←W(l)h~(l−1)\tilde{z}_{\text{pre}}^{(l)} \leftarrow W^{(l)} \tilde{h}^{(l-1)}
        z~(l)←batchnorm(z~pre(l))+n(l)\tilde{z}^{(l)} \leftarrow \text{batchnorm}(\tilde{z}_{\text{pre}}^{(l)}) + n^{(l)}
        h~(l)←ϕ(γ(l)⊙(z~(l)+β(l)))\tilde{h}^{(l)} \leftarrow \phi(\gamma^{(l)} \odot (\tilde{z}^{(l)} + \beta^{(l)}))
    end for
    P(y~∣x)←h~(L)P(\tilde{y} \mid x) \leftarrow \tilde{h}^{(L)}
    # 2. Clean Encoder Pass (computes denoising targets and clean prediction)
    h(0)←z(0)←x(n)h^{(0)} \leftarrow z^{(0)} \leftarrow x(n)
    for l=1l = 1 to LL do
        zpre(l)←W(l)h(l−1)z_{\text{pre}}^{(l)} \leftarrow W^{(l)} h^{(l-1)}
        μ(l)←batchmean(zpre(l))\mu^{(l)} \leftarrow \text{batchmean}(z_{\text{pre}}^{(l)})
        σ(l)←batchstd(zpre(l))\sigma^{(l)} \leftarrow \text{batchstd}(z_{\text{pre}}^{(l)})
        z(l)←batchnorm(zpre(l))z^{(l)} \leftarrow \text{batchnorm}(z_{\text{pre}}^{(l)})
        h(l)←ϕ(γ(l)⊙(z(l)+β(l)))h^{(l)} \leftarrow \phi(\gamma^{(l)} \odot (z^{(l)} + \beta^{(l)}))
    end for
    P(y∣x)←h(L)P(y \mid x) \leftarrow h^{(L)}
    # 3. Decoder Denoising Pass
    for l=Ll = L down to 00 do
        if l==Ll == L then
            u(L)←batchnorm(h~(L))u^{(L)} \leftarrow \text{batchnorm}(\tilde{h}^{(L)})
        else
            u(l)←batchnorm(V(l+1)z^(l+1))u^{(l)} \leftarrow \text{batchnorm}(V^{(l+1)} \hat{z}^{(l+1)})
        end if
        for each neuron index ii do
            z^i(l)←gi(z~i(l),ui(l))\hat{z}_i^{(l)} \leftarrow g_i(\tilde{z}_i^{(l)}, u_i^{(l)})
            z^i,BN(l)←(z^i(l)−μi(l))/σi(l)\hat{z}_{i,\text{BN}}^{(l)} \leftarrow (\hat{z}_i^{(l)} - \mu_i^{(l)}) / \sigma_i^{(l)}
        end for
    end for
    # 4. Total Cost Computation
    C←0C \leftarrow 0
    if t(n)t(n) is provided then
        C←−log⁡P(y~=t(n)∣x(n))C \leftarrow -\log P(\tilde{y} = t(n) \mid x(n))
    end if
    C←C+∑l=0Lλl∥z(l)−z^BN(l)∥2C \leftarrow C + \sum_{l=0}^L \lambda_l \| z^{(l)} - \hat{z}_{\text{BN}}^{(l)} \|^2
    return P(y∣x)P(y \mid x), CC

    During inference on test data, the output prediction is computed using the clean feedforward path P(y∣x)=h(L)P(y \mid x) = h^{(L)} without noise injection or decoder execution.

  3. Knowl 3 — Batch-Normalized Layer-Wise Denoising Objective

    equation

    The unsupervised loss CdC_d of the Ladder network is the weighted sum of mean squared reconstruction errors across all L+1L+1 layers (from input layer l=0l=0 to top layer l=Ll=L):

    Cd=∑l=0LλlCd(l)=∑l=0LλlNml∑n=1N∥z(l)(n)−z^BN(l)(n)∥2C_d = \sum_{l=0}^L \lambda_l C_d^{(l)} = \sum_{l=0}^L \frac{\lambda_l}{N m_l} \sum_{n=1}^N \| z^{(l)}(n) - \hat{z}_{\text{BN}}^{(l)}(n) \|^2

    where NN is the number of training samples in the minibatch, mlm_l is the width (dimension) of layer ll, λl\lambda_l is a non-negative layer-specific cost multiplier, z(l)(n)z^{(l)}(n) is the clean batch-normalized latent representation at layer ll, and z^BN(l)(n)\hat{z}_{\text{BN}}^{(l)}(n) is the batch-normalized decoder reconstruction:

    z^BN(l)=z^(l)−μ(l)σ(l)\hat{z}_{\text{BN}}^{(l)} = \frac{\hat{z}^{(l)} - \mu^{(l)}}{\sigma^{(l)}}

    Here, μ(l)\mu^{(l)} and σ(l)\sigma^{(l)} denote the batch mean and batch standard deviation of the clean preactivations zpre(l)z_{\text{pre}}^{(l)} from the clean encoder pass. Normalizing the reconstruction target to the batch-normalized coordinate frame prevents the network from exploiting minibatch-correlated batch normalization noise between clean z(l)z^{(l)} and corrupted z~(l)\tilde{z}^{(l)} to achieve trivial copy solutions z^(l)≈z~(l)\hat{z}^{(l)} \approx \tilde{z}^{(l)}.

  4. Knowl 4 — Denoising Function Parametrization for Conditional Gaussian Latents

    equation

    Assuming the latent representation z(l)z^{(l)} is conditionally Gaussian given the representation of the layer above z(l+1)z^{(l+1)}, the optimal Bayes reconstruction z^i(l)\hat{z}_i^{(l)} for neuron ii at layer ll from its corrupted value z~i(l)\tilde{z}_i^{(l)} and top-down projection ui(l)u_i^{(l)} is parameterized as:

    z^i(l)=gi(z~i(l),ui(l))=(z~i(l)−μi(ui(l)))νi(ui(l))+μi(ui(l))\hat{z}_i^{(l)} = g_i(\tilde{z}_i^{(l)}, u_i^{(l)}) = \left(\tilde{z}_i^{(l)} - \mu_i(u_i^{(l)})\right) \nu_i(u_i^{(l)}) + \mu_i(u_i^{(l)})

    where u(l)=NB(V(l+1)z^(l+1))u^{(l)} = \text{NB}(V^{(l+1)} \hat{z}^{(l+1)}) is a batch-normalized linear projection of the higher-level reconstruction z^(l+1)\hat{z}^{(l+1)}, and μi(u)\mu_i(u) and νi(u)\nu_i(u) are scalar nonlinear functions parameterized with trainable weights a1,i(l),…,a10,i(l)a_{1,i}^{(l)}, \dots, a_{10,i}^{(l)}:

    μi(u)=a1,i(l)sigmoid(a2,i(l)u+a3,i(l))+a4,i(l)u+a5,i(l)\mu_i(u) = a_{1,i}^{(l)} \text{sigmoid}\left(a_{2,i}^{(l)} u + a_{3,i}^{(l)}\right) + a_{4,i}^{(l)} u + a_{5,i}^{(l)}

    νi(u)=a6,i(l)sigmoid(a7,i(l)u+a8,i(l))+a9,i(l)u+a10,i(l)\nu_i(u) = a_{6,i}^{(l)} \text{sigmoid}\left(a_{7,i}^{(l)} u + a_{8,i}^{(l)}\right) + a_{9,i}^{(l)} u + a_{10,i}^{(l)}

    Conditioned on ui(l)u_i^{(l)}, the denoising function is linear with respect to z~i(l)\tilde{z}_i^{(l)}, where μi(ui(l))\mu_i(u_i^{(l)}) estimates the prior mean and νi(ui(l))\nu_i(u_i^{(l)}) modulates the weighting between the corrupted observation and the prior mean based on noise and signal variances.

  5. Knowl 5 — The Gamma-Model Architecture Variant

    model/method

    The Γ\Gamma-model is a simplified variant of the Ladder network obtained by setting all layer-wise unsupervised cost multipliers λl=0\lambda_l = 0 for l<Ll < L, leaving only the denoising cost on the topmost layer λL>0\lambda_L > 0.

    In this configuration, the intermediate decoder layers and lateral connections are omitted. The architecture retains both the clean feedforward pathway and the corrupted feedforward pathway of the encoder. Denoising regularization is applied solely at the output layer LL, penalizing discrepancies between the predictions of the corrupted and clean pathways. This formulation can be applied directly to standard feedforward multi-layer perceptrons or convolutional neural networks without constructing an inverted decoder network.

  6. Knowl 6 — Ladder Network Adaptation for Convolutional Architectures

    model/method

    Adapting the Ladder network to convolutional neural networks (CNNs) involves specific modifications to the encoder and decoder structures:

    1. Mirrored Convolutions and Parameter Sharing: Decoder layers mirror encoder convolution operations with transposed convolutions (deconvolutions). Weight sharing principles are applied to the denoising function gg across spatial locations of each feature map.

    2. Pooling and Depooling: Strided pooling operations in the encoder are treated as explicit layers with independent batch normalization. In the decoder, downsampling from encoder pooling is inverted via upsampling by spatial copying. This provides distinct reconstruction targets at each pooling stage, assisting the decoder in recovering spatial information lost during downsampling.

    3. Batch Normalization: Per-channel batch normalization is applied across all layers, including pooling stages, to stabilize training and facilitate layer-wise denoising targets.

  7. Knowl 7 — Permutation-Invariant MNIST Semi-Supervised Classification Benchmark

    data/table

    Test error rates on the permutation-invariant MNIST dataset across different numbers of labeled examples (N=100N = 100, N=1000N = 1000, and all 60,00060{,}000 labels), comparing the Ladder network against baseline and previous semi-supervised methods. Evaluated on a 784-1000-500-250-250-250-10 fully connected architecture trained using Adam optimization.

    Method 100 labels 1000 labels All labels
    Semi-supervised Embedding (Weston et al., 2012) 16.86% 5.73% 1.50%
    Transductive SVM (Weston et al., 2012) 16.81% 5.38% 1.40%
    MTC (Rifai et al., 2011) 12.03% 3.64% 0.81%
    Pseudo-label (Lee, 2013) 10.49% 3.46% –
    AtlasRBF (Pitelis et al., 2014) 8.10% (±0.95\pm 0.95) 3.68% (±0.12\pm 0.12) 1.31%
    DGN (Kingma et al., 2014) 3.33% (±0.14\pm 0.14) 2.40% (±0.02\pm 0.02) 0.96%
    DBM, Dropout (Srivastava et al., 2014) – – 0.79%
    Adversarial (Goodfellow et al., 2015) – – 0.78%
    Virtual Adversarial (Miyato et al., 2015) 2.12% 1.32% 0.64% (±0.03\pm 0.03)
    Baseline: MLP, BN, Gaussian noise 21.74% (±1.77\pm 1.77) 5.70% (±0.20\pm 0.20) 0.80% (±0.03\pm 0.03)
    Γ\Gamma-model (top-level cost only) 3.06% (±1.44\pm 1.44) 1.53% (±0.10\pm 0.10) 0.78% (±0.03\pm 0.03)
    Ladder, bottom-level cost only 1.09% (±0.32\pm 0.32) 0.90% (±0.05\pm 0.05) 0.59% (±0.03\pm 0.03)
    Ladder, full 1.06% (±0.37\pm 0.37) 0.84% (±0.08\pm 0.08) 0.57% (±0.02\pm 0.02)

    The full Ladder network reduces test error on 100 labeled examples to 1.06%1.06\%, outperforming prior methods by a substantial margin, and achieves a state-of-the-art error of 0.57%0.57\% in the fully labeled permutation-invariant setting.

  8. Knowl 8 — Convolutional Ladder and Gamma-Model Results on MNIST and CIFAR-10

    data/table

    Semi-supervised performance of convolutional implementations of the Ladder network and Γ\Gamma-model on standard MNIST and CIFAR-10 without data augmentation.

    Dataset Model Labeled Samples Test Error
    MNIST (CNN) EmbedCNN (Weston et al., 2012) 100 7.75%
    MNIST (CNN) SWWAE (Zhao et al., 2015) 100 9.17%
    MNIST (CNN) Baseline: Conv-Small (supervised only) 100 6.43% (±0.84\pm 0.84)
    MNIST (CNN) Conv-FC (Full Ladder, 1st layer conv) 100 0.99% (±0.15\pm 0.15)
    MNIST (CNN) Conv-Small, Γ\Gamma-model 100 0.89% (±0.50\pm 0.50)
    CIFAR-10 Spike-and-Slab Sparse Coding (Goodfellow et al., 2012) 4000 31.90%
    CIFAR-10 Baseline: Conv-Large (supervised only) 4000 23.33% (±0.61\pm 0.61)
    CIFAR-10 Conv-Large, Γ\Gamma-model 4000 20.40% (±0.47\pm 0.47)

    On MNIST with 100 labels, the convolutional Γ\Gamma-model achieves 0.89%0.89\% test error and Conv-FC achieves 0.99%0.99\%, outperforming SWWAE (9.17%9.17\%). On CIFAR-10 with 4000 labels, the Γ\Gamma-model reduces error from the supervised baseline of 23.33%23.33\% to 20.40%20.40\%.

  9. Knowl 9 — Ablation of Lateral Modulation in the Denoising Function

    data/table

    Ablation experiment evaluating alternative functional forms of the denoising function g(z~,u)g(\tilde{z}, u) on semi-supervised MNIST with 100 and 1000 labeled samples (with standard deviation of corruption noise fixed at 0.3).

    Parametrization of g(z~,u)g(\tilde{z}, u) 100 labels 1000 labels
    Proposed gg: Gaussian zz (linear in z~\tilde{z} with nonlinear μ(u),ν(u)\mu(u), \nu(u)) 1.06% (±0.07\pm 0.07) 1.03% (±0.06\pm 0.06)
    g1g_1: Miniature MLP with z~u\tilde{z}u interaction term 1.11% (±0.07\pm 0.07) 1.11% (±0.06\pm 0.06)
    g2g_2: No augmented interaction term z~u\tilde{z}u 2.03% (±0.09\pm 0.09) 1.70% (±0.08\pm 0.08)
    g3g_3: Linear gg with z~u\tilde{z}u interaction term 1.49% (±0.10\pm 0.10) 1.30% (±0.08\pm 0.08)
    g4g_4: Modulation affects only the mean of p(z∣u)p(z \mid u) 2.90% (±1.19\pm 1.19) 2.11% (±0.45\pm 0.45)

    The results demonstrate that multiplicative modulation of the corrupted lateral feed z~\tilde{z} by the top-down signal uu (present in the proposed gg, g1g_1, and g3g_3) is critical; omitting the interaction term (g2g_2) or restricting uu to purely additive mean shifts (g4g_4) substantially degrades classification performance.

Coverage note — None was omitted; all contributed models, equations, algorithms, empirical benchmarks, and ablation studies are covered.

References

  1. 1.Bastien, F., Lamblin, P., Pascanu, R., Bergstra, J., Goodfellow, I. J., Bergeron, A., Bouchard, N., and Bengio, Y. (2012). Theano: new features and speed improvements. Deep Learning and Unsupervised Feature Learning NIPS 2012 Workshop.
  2. 2.Bengio, Y. (2014). How auto-encoders could provide credit assignment in deep networks via target propagation. arXiv:1407.7906.
  3. 3.Bengio, Y., Yao, L., Alain, G., and Vincent, P. (2013). Generalized denoising auto-encoders as generative models. In Advances in Neural Information Processing Systems 26 (NIPS 2013), pages 899–907.
  4. 4.Bergstra, J., Breuleux, O., Bastien, F., Lamblin, P., Pascanu, R., Desjardins, G., Turian, J., Warde-Farley, D., and Bengio, Y. (2010). Theano: a CPU and GPU math expression compiler. In Proceedings of the Python for Scientific Computing Conference (SciPy 2010). Oral Presentation.
  5. 5.Blum, A. and Mitchell, T. (1998). Combining labeled and unlabeled data with co-training. In Proc. of the Eleventh Annual Conference on Computational Learning Theory (COLT ’98), pages 92–100.
  6. 6.Chapelle, O., Schölkopf, B., Zien, A., et al. (2006). Semi-supervised learning. MIT Press.
  7. 7.Dosovitskiy, A., Springenberg, J. T., Riedmiller, M., and Brox, T. (2014). Discriminative unsupervised feature learning with convolutional neural networks. In Advances in Neural Information Processing Systems 27 (NIPS 2014), pages 766–774.
  8. 8.Goodfellow, I., Bengio, Y., and Courville, A. C. (2012). Large-scale feature learning with spike-and-slab sparse coding. In Proc. of ICML 2012, pages 1439–1446.
  9. 9.Goodfellow, I., Mirza, M., Courville, A., and Bengio, Y. (2013a). Multi-prediction deep Boltzmann machines. In Advances in Neural Information Processing Systems 26 (NIPS 2013), pages 548–556.
  10. 10.Goodfellow, I., Shlens, J., and Szegedy, C. (2015). Explaining and harnessing adversarial examples. In the International Conference on Learning Representations (ICLR 2015). arXiv:1412.6572.
  11. 11.Goodfellow, I. J., Warde-Farley, D., Mirza, M., Courville, A., and Bengio, Y. (2013b). Maxout networks. In Proc. of ICML 2013.
  12. 12.Gregor, K., Danihelka, I., Mnih, A., Blundell, C., and Wierstra, D. (2014). Deep autoregressive networks. In Proc. of ICML 2014, Beijing, China.
  13. 13.Hinton, G. E. and Salakhutdinov, R. R. (2006). Reducing the dimensionality of data with neural networks. Science, 313(5786), 504–507.
  14. 14.Ioffe, S. and Szegedy, C. (2015). Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International Conference on Machine Learning (ICML), pages 448–456.
  15. 15.Kingma, D. and Ba, J. (2015). Adam: A method for stochastic optimization. In the International Conference on Learning Representations (ICLR 2015), San Diego. arXiv:1412.6980.
  16. 16.Kingma, D. P., Mohamed, S., Rezende, D. J., and Welling, M. (2014). Semi-supervised learning with deep generative models. In Advances in Neural Information Processing Systems 27 (NIPS 2014), pages 3581–3589.
  17. 17.Lee, D.-H. (2013). Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks. In Workshop on Challenges in Representation Learning, ICML 2013.
  18. 18.McLachlan, G. (1975). Iterative reclassification procedure for constructing an asymptotically optimal rule of allocation in discriminant analysis. J. American Statistical Association, 70, 365–369.
  19. 19.Miyato, T., Maeda, S., Koyama, M., Nakae, K., and Ishii, S. (2015). Distributional smoothing by virtual adversarial examples. arXiv:1507.00677.
  20. 20.Pezeshki, M., Fan, L., Brakel, P., Courville, A., and Bengio, Y. (2015). Deconstructing the ladder network architecture. arXiv:1511.06430.
  21. 21.Pitelis, N., Russell, C., and Agapito, L. (2014). Semi-supervised learning using an unsupervised atlas. In Machine Learning and Knowledge Discovery in Databases (ECML PKDD 2014), pages 565–580. Springer.
  22. 22.Raiko, T., Berglund, M., Alain, G., and Dinh, L. (2015). Techniques for learning binary stochastic feedforward neural networks. In ICLR 2015, San Diego.
  23. 23.Ranzato, M. A. and Szummer, M. (2008). Semi-supervised learning of compact document representations with deep networks. In Proc. of ICML 2008, pages 792–799. ACM.
  24. 24.Rasmus, A., Raiko, T., and Valpola, H. (2015a). Denoising autoencoder with modulated lateral connections learns invariant representations of natural images. arXiv:1412.7210.
  25. 25.Rasmus, A., Valpola, H., and Raiko, T. (2015b). Lateral connections in denoising autoencoders support supervised learning. arXiv:1504.08215.
  26. 26.Rifai, S., Dauphin, Y. N., Vincent, P., Bengio, Y., and Muller, X. (2011). The manifold tangent classifier. In Advances in Neural Information Processing Systems 24 (NIPS 2011), pages 2294–2302.
  27. 27.Särelä, J. and Valpola, H. (2005). Denoising source separation. JMLR, 6, 233–272.
  28. 28.Sietsma, J. and Dow, R. J. (1991). Creating artificial neural networks that generalize. Neural networks, 4(1), 67–79.
  29. 29.Springenberg, J. T., Dosovitskiy, A., Brox, T., and Riedmiller, M. A. (2014). Striving for simplicity: The all convolutional net. arxiv:1412.6806.
  30. 30.Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. (2014). Dropout: A simple way to prevent neural networks from overfitting. JMLR, 15(1), 1929–1958.
  31. 31.Suddarth, S. C. and Kergosien, Y. (1990). Rule-injection hints as a means of improving network performance and learning time. In Proceedings of the EURASIP Workshop 1990 on Neural Networks, pages 120–129. Springer.
  32. 32.Szummer, M. and Jaakkola, T. (2003). Partially labeled classification with Markov random walks. Advances in Neural Information Processing Systems 15 (NIPS 2002), 14, 945–952.
  33. 33.Titterington, D., Smith, A., and Makov, U. (1985). Statistical analysis of finite mixture distributions. In Wiley Series in Probability and Mathematical Statistics. Wiley.
  34. 34.Valpola, H. (2015). From neural PCA to deep unsupervised learning. In Adv. in Independent Component Analysis and Learning Machines, pages 143–171. Elsevier. arXiv:1411.7783.
  35. 35.van Merriënboer, B., Bahdanau, D., Dumoulin, V., Serdyuk, D., Warde-Farley, D., Chorowski, J., and Bengio, Y. (2015). Blocks and fuel: Frameworks for deep learning. CoRR, abs/1506.00619.
  36. 36.Vincent, P., Larochelle, H., Lajoie, I., Bengio, Y., and Manzagol, P.-A. (2010). Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion. JMLR, 11, 3371–3408.
  37. 37.Weston, J., Ratle, F., Mobahi, H., and Collobert, R. (2012). Deep learning via semi-supervised embedding. In Neural Networks: Tricks of the Trade, pages 639–655. Springer.
  38. 38.Zeiler, M. D., Taylor, G. W., and Fergus, R. (2011). Adaptive deconvolutional networks for mid and high level feature learning. In ICCV 2011, pages 2018–2025. IEEE.
  39. 39.Zhang, T. and Oles, F. (2000). The value of unlabeled data for classification problems. In Proc. of ICML 2000, pages 1191–1198.
  40. 40.Zhao, J., Mathieu, M., Goroshin, R., and Lecun, Y. (2015). Stacked what-where auto-encoders. arXiv:1506.02351.

Citation

MLA
Rasmus, A., et al. “Semi-Supervised Learning with Ladder Networks”. arXiv, 2015, http://arxiv.org/abs/1507.02672v2.
APA
Rasmus, A., Valpola, H., Honkala, M., Berglund, M., & Raiko, T. (2015). Semi-Supervised Learning with Ladder Networks. arXiv. http://arxiv.org/abs/1507.02672v2
Chicago
Rasmus, A., H. Valpola, M. Honkala, M. Berglund, and T. Raiko. 2015. “Semi-Supervised Learning with Ladder Networks”. arXiv. http://arxiv.org/abs/1507.02672v2.
Harvard
Rasmus, A. et al. (2015) “Semi-Supervised Learning with Ladder Networks”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1507.02672v2.
Vancouver
1. Rasmus A, Valpola H, Honkala M, Berglund M, Raiko T (2015) Semi-Supervised Learning with Ladder Networks. arXiv

BibTeX

@article{rasmus2015semi,
  title = {Semi-Supervised Learning with Ladder Networks},
  author = {Rasmus, Antti and Valpola, Harri and Honkala, Mikko and Berglund, Mathias and Raiko, Tapani},
  year = {2015},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1507.02672v2},
  eprint = {1507.02672}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors