Contractive Auto-Encoders: Explicit Invariance During Feature Extraction

Salah RifaiP. VincentX. MullerXavier GlorotYoshua Bengio

article2011ICML1,588 citations

Proposes penalizing the Frobenius norm of the encoder's Jacobian matrix during auto-encoder training to encourage localized invariance to small input perturbations and extract features that improve downstream classification.

Listen

Building accurate machine learning systems often requires extracting meaningful representations from unlabeled data. A central question is how to design automated feature extractors that capture genuine patterns while remaining unaffected by irrelevant noise and small variations. The article introduces and evaluates the Contractive Auto-Encoder, an unsupervised feature-learning method that incorporates a mathematical penalty into standard auto-encoder training to enforce localized stability and extract robust data representations.

The authors conducted comparative empirical evaluations across standard image benchmark datasets, including MNIST, a grayscale version of CIFAR-10, and seven variations featuring complex visual transformations. The method adds an analytic penalty term—specifically measuring how much the encoder's outputs change in response to tiny shifts in the input—directly to the standard reconstruction objective. The resulting model was evaluated as an unsupervised pre-training step for multi-layer neural networks, comparing its final classification performance and geometric properties against baseline auto-encoders, weight-decay regularized auto-encoders, denoising auto-encoders, and Restricted Boltzmann Machines.

The experiments produced several key findings. First, the Contractive Auto-Encoder achieved the lowest classification error among single-layer models on standard benchmarks, recording a 1.14% error rate on MNIST and 47.86% on grayscale CIFAR-10. Second, the degree of local contraction directly correlated with superior downstream classification performance. Third, geometric analyses showed that the proposed penalty effectively contracts directions unrelated to the data while preserving the essential directions of variation necessary for faithful reconstruction. Finally, stacking these models into deep architectures proved highly effective: two-layer contractive models frequently matched or outperformed established three-layer alternative networks across challenging image tasks, such as achieving a 2.48% error rate on standard digit subsets and 1.21% on synthetic shape recognition.

These results demonstrate that an analytic penalty on feature sensitivity provides a principled and computationally efficient way to achieve invariance to noise without relying on randomized corruption techniques. The model automatically balances reconstruction accuracy with local stability, capturing the underlying low-dimensional structure of complex data. Practitioners seeking to improve representation learning and semi-supervised pipelines can adopt contractive auto-encoders to enhance model accuracy with computational costs that remain comparable to standard auto-encoders. Future implementations can explore deeper stacked configurations or test the approach on broader domains beyond visual benchmarks.

Cover for Contractive Auto-Encoders: Explicit Invariance During Feature Extraction

Abstract

We present in this paper a novel approach for training deterministic auto-encoders. We show that by adding a well chosen penalty term to the classical reconstruction cost function, we can achieve results that equal or surpass those attained by other regularized auto-encoders as well as denoising auto-encoders on a range of datasets. This penalty term corresponds to the Frobenius norm of the Jacobian matrix of the encoder activations with respect to the input. We show that this penalty term results in a localized space contraction which in turn yields robust features on the activation layer. Furthermore, we show how this penalty term is related to both regularized auto-encoders and denoising auto-encoders and how it can be seen as a link between deterministic and non-deterministic auto-encoders. We find empirically that this penalty helps to carve a representation that better captures the local directions of variation dictated by the data, corresponding to a lower-dimensional non-linear manifold, while being more invariant to the vast majority of directions orthogonal to the manifold. Finally, we show that by using the learned features to initialize a MLP, we achieve state of the art classification error on a range of datasets, surpassing other methods of pretraining.

Table of Contents

  • 1. Introduction
  • 2. How to extract robust features
  • 3. Auto-encoders variants
  • 4. Contractive auto-encoders (CAE)
  • Computational considerations
  • 5. Experiments and results
  • 5.1. Classification performance
  • 5.2. Closer examination of the contraction
  • 5.3. Discussion: Local Space Contraction
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — Contractive Auto-Encoder (CAE) Objective Function

    model/method

    The Contractive Auto-Encoder (CAE) is an unsupervised feature extraction model designed to learn representations that are robust to small perturbations around the training points. Given a dataset of training examples Dn={x(1),…,x(n)}\mathcal{D}_n = \{x^{(1)}, \dots, x^{(n)}\} where x∈Rdxx \in \mathbb{R}^{d_x}, an encoder f:Rdx→Rdhf: \mathbb{R}^{d_x} \to \mathbb{R}^{d_h}, and a decoder g:Rdh→Rdxg: \mathbb{R}^{d_h} \to \mathbb{R}^{d_x}, the CAE objective minimizes the sum of a standard reconstruction error L(x,g(f(x)))L(x, g(f(x))) and a regularization term penalizing the sensitivity of the representation with respect to the input, weighted by a hyperparameter λ>0\lambda > 0:

    JCAE(θ)=∑x∈Dn(L(x,g(f(x)))+λ∥Jf(x)∥F2)J_{\mathrm{CAE}}(\theta) = \sum_{x \in \mathcal{D}_n} \left( L(x, g(f(x))) + \lambda \|J_f(x)\|_F^2 \right)

    where θ={W,bh,by}\theta = \{W, b_h, b_y\} denotes the model parameters, and ∥Jf(x)∥F2\|J_f(x)\|_F^2 is the squared Frobenius norm of the Jacobian matrix Jf(x)J_f(x) of the encoder mapping ff evaluated at input xx:

    ∥Jf(x)∥F2=∑i=1dh∑j=1dx(∂hi(x)∂xj)2\|J_f(x)\|_F^2 = \sum_{i=1}^{d_h} \sum_{j=1}^{d_x} \left( \frac{\partial h_i(x)}{\partial x_j} \right)^2

    with h=f(x)=sf(Wx+bh)h = f(x) = s_f(W x + b_h) being the hidden activation vector, W∈Rdh×dxW \in \mathbb{R}^{d_h \times d_x} the encoder weight matrix, bh∈Rdhb_h \in \mathbb{R}^{d_h} the encoder bias, and sfs_f a non-linear activation function (such as the logistic sigmoid sf(z)=11+e−zs_f(z) = \frac{1}{1 + e^{-z}}). The decoder reconstructs y=g(h)=sg(WTh+by)y = g(h) = s_g(W^T h + b_y) using tied weights WTW^T, decoder bias by∈Rdxb_y \in \mathbb{R}^{d_x}, and activation sgs_g.

  2. Knowl 2 — Analytic Jacobian Penalty for Sigmoid Encoders and Computational Complexity

    equation

    When the encoder mapping f(x)f(x) employs an element-wise logistic sigmoid non-linearity sf(z)=11+e−zs_f(z) = \frac{1}{1 + e^{-z}} on the affine transformation a(x)=Wx+bha(x) = W x + b_h, the squared Frobenius norm of the encoder Jacobian Jf(x)∈Rdh×dxJ_f(x) \in \mathbb{R}^{d_h \times d_x} simplifies to an exact analytic expression:

    ∥Jf(x)∥F2=∑i=1dh(hi(1−hi))2∑j=1dxWij2\|J_f(x)\|_F^2 = \sum_{i=1}^{d_h} \left( h_i (1 - h_i) \right)^2 \sum_{j=1}^{d_x} W_{ij}^2

    where hi=sf(Wi⋅x+bh,i)h_i = s_f(W_{i\cdot} x + b_{h,i}) is the activation of the ii-th hidden unit, and WijW_{ij} is the weight connecting input dimension jj to hidden unit ii.

    Evaluating this penalty and its gradient with respect to the parameters θ={W,bh}\theta = \{W, b_h\} requires O(dx×dh)\mathcal{O}(d_x \times d_h) operations per training example, which is identical in asymptotic computational complexity to computing the reconstruction cost and its gradient.

  3. Knowl 3 — Geometric Mechanism of Manifold Learning via Local Space Contraction

    theoretical result

    The regularized objective of the Contractive Auto-Encoder balances two opposing forces:

    1. Contraction force: The Jacobian penalty ∥Jf(x)∥F2\|J_f(x)\|_F^2 exerts a contracting pressure that encourages the representation f(x)f(x) to be locally constant in all directions in input space around each training sample xx.
    2. Reconstruction force: The reconstruction cost L(x,g(f(x)))L(x, g(f(x))) penalizes information loss, requiring the representation h=f(x)h = f(x) to retain enough information to distinguish and reconstruct distinct training examples.

    Under the assumption that the data distribution concentrates near a low-dimensional manifold, variations present in the training set correspond to local tangent directions along the manifold, whereas small or rare variations correspond to orthogonal directions. Because preserving variations along the manifold is strictly required to reconstruct neighboring data points, the encoder resists contraction only along tangent directions. In contrast, orthogonal directions (where data density falls off sharply) are heavily contracted (yielding small singular values of Jf(x)J_f(x)). Consequently, the representation achieves local invariance in the directions orthogonal to the underlying data manifold while capturing variations along the manifold.

  4. Knowl 4 — Theoretical Comparison Between Contractive and Denoising Auto-Encoders

    theoretical result

    Contractive Auto-Encoders (CAE) and Denoising Auto-Encoders (DAE) achieve robustness through distinct mathematical mechanisms:

    1. Target of Robustness: The CAE explicitly penalizes the sensitivity of the representation f(x)f(x) through ∥Jf(x)∥F2\|J_f(x)\|_F^2. The DAE encourages robustness of the entire reconstruction (g∘f)(x)(g \circ f)(x) across stochastically corrupted inputs x~∼q(x~∣x)\tilde{x} \sim q(\tilde{x}|x), which only indirectly encourages invariance of the latent features because the invariance burden is shared between encoder ff and decoder gg.
    2. Stochastic vs. Analytic Formulation: The DAE optimizes an empirical expectation over sampled corruptions: JDAE(θ)=∑x∈DnEx~∼q(x~∣x)[L(x,g(f(x~)))]J_{\mathrm{DAE}}(\theta) = \sum_{x \in \mathcal{D}_n} \mathbb{E}_{\tilde{x} \sim q(\tilde{x}|x)} [L(x, g(f(\tilde{x})))] In the asymptotic limit of infinitesimal isotropic Gaussian noise x~=x+ϵ\tilde{x} = x + \epsilon with ϵ∼N(0,σ2I)\epsilon \sim \mathcal{N}(0, \sigma^2 I), Taylor expansion reveals that the DAE objective analytically approximates a penalty on the squared Frobenius norm of the reconstruction Jacobian, ∥Jg∘f(x)∥F2\|J_{g \circ f}(x)\|_F^2, in contrast to the CAE penalty on the encoder Jacobian ∥Jf(x)∥F2\|J_f(x)\|_F^2 evaluated deterministically at the clean training points.
  5. Knowl 5 — Relationship Between the Contractive Penalty, Weight Decay, and Sparsity

    theoretical result

    The contractive penalty ∥Jf(x)∥F2\|J_f(x)\|_F^2 relates directly to classical regularization mechanisms:

    • Weight Decay: For a linear encoder where sfs_f is the identity function, the Jacobian is Jf(x)=WJ_f(x) = W, and its squared Frobenius norm is ∥Jf(x)∥F2=∑ijWij2\|J_f(x)\|_F^2 = \sum_{ij} W_{ij}^2. Under a linear encoder, the CAE objective is algebraically identical to standard L2L_2 weight decay (JAE+wdJ_{\mathrm{AE+wd}}). With a non-linear sigmoid activation, contraction can be achieved either by keeping weights WW small or by driving hidden units hih_i into their saturated regimes (hi→0h_i \to 0 or hi→1h_i \to 1), where hi(1−hi)→0h_i(1-h_i) \to 0.
    • Sparse Auto-Encoders: Sparse auto-encoders enforce hidden activations to remain near zero for most units on any given sample. Because the left tail of the logistic sigmoid is asymptotically flat with near-zero first derivatives, sparse representations implicitly yield small entries in Jf(x)J_f(x), thereby producing an implicit contractive mapping despite not explicitly penalizing derivatives in their loss function.
  6. Knowl 6 — Isotropic Contraction Ratio and Contraction Curves

    definition

    The isotropic contraction ratio quantifies the non-local contraction of an encoder mapping f:Rdx→Rdhf: \mathbb{R}^{d_x} \to \mathbb{R}^{d_h} as a function of perturbation radius rr. For a reference point x0∈Rdxx_0 \in \mathbb{R}^{d_x} sampled from a dataset and a perturbed point x1x_1 sampled uniformly at random on a hypersphere of radius rr centered at x0x_0 (∥x1−x0∥=r\|x_1 - x_0\| = r), the isotropic contraction ratio is defined as:

    Contraction Ratio(r)=Ex0,x1[∥f(x1)−f(x0)∥∥x1−x0∥]\text{Contraction Ratio}(r) = \mathbb{E}_{x_0, x_1} \left[ \frac{\|f(x_1) - f(x_0)\|}{\|x_1 - x_0\|} \right]

    In the limit r→0r \to 0, this ratio converges to the directional derivative (the magnitude of the Jacobian mapping). Plotting this expectation across varying radii rr produces a contraction curve, characterizing how representation sensitivity transitions from immediate neighborhoods of data points to distant regions of the input space.

  7. Knowl 7 — Manifold Dimensionality Estimation via Jacobian Singular Value Spectrum

    empirical result

    The singular value decomposition (SVD) of the encoder Jacobian matrix Jf(x)J_f(x) reveals the directional sensitivity of the learned representation at input xx. Singular values correspond to the rate of variation of the representation along the singular vector directions:

    • Large singular values identify the local tangent directions along which the model preserves variations in the data.
    • Small singular values identify orthogonal directions where the input space is strongly contracted.

    On the CIFAR-bw dataset, the singular value spectrum of the CAE exhibits a substantially sharper drop than standard auto-encoders (AE), weight-decay regularized auto-encoders (AE+wd), and denoising auto-encoders (DAE). The CAE concentrates its sensitivity in a significantly smaller number of large singular values, demonstrating that it restricts variance to a lower-dimensional manifold while maintaining invariance across the remaining input dimensions.

  8. Knowl 8 — Classification Performance and Representation Contraction on MNIST and CIFAR-bw

    data/table

    Unsupervised pre-training of a single-layer neural network with 1000 hidden units using the Contractive Auto-Encoder (CAE) achieves lower classification error after supervised fine-tuning compared to basic auto-encoders (AE), weight-decay regularized auto-encoders (AE+wd), binary Restricted Boltzmann Machines (RBM-binary), and Denoising Auto-Encoders with Gaussian (DAE-g) or binary masking noise (DAE-b).

    Lower test classification error strongly correlates with a lower average Frobenius norm of the encoder Jacobian ∥Jf(x)∥F\|J_f(x)\|_F and a higher unit saturation rate SAT\text{SAT} (defined as the average percentage of hidden units with activations <0.05< 0.05 or >0.95> 0.95 per example):

    Model Test Error (%) Average ∥Jf(x)∥F\|J_f(x)\|_F SAT (%)
    MNIST
    CAE 1.14 0.73×10−40.73 \times 10^{-4} 86.36
    DAE-g 1.18 0.86×10−40.86 \times 10^{-4} 17.77
    RBM-binary 1.30 2.50×10−42.50 \times 10^{-4} 78.59
    DAE-b 1.57 7.87×10−47.87 \times 10^{-4} 68.19
    AE+wd 1.68 5.00×10−45.00 \times 10^{-4} 12.97
    AE 1.78 17.5×10−417.5 \times 10^{-4} 49.90
    CIFAR-bw
    CAE 47.86 2.40×10−52.40 \times 10^{-5} 85.65
    DAE-b 49.03 4.85×10−54.85 \times 10^{-5} 80.66
    DAE-g 54.81 4.94×10−54.94 \times 10^{-5} 19.90
    AE+wd 55.03 34.9×10−534.9 \times 10^{-5} 23.04
    AE 55.47 44.9×10−544.9 \times 10^{-5} 22.57
  9. Knowl 9 — Benchmark Performance of Stacked Contractive Auto-Encoders on MNIST Variations

    data/table

    Stacked Contractive Auto-Encoders with 1 and 2 layers (CAE-1 and CAE-2) were evaluated on the MNIST variations benchmark suite (each dataset with 10,000 training, 2,000 validation, and 50,000 test examples) and compared against an RBF Support Vector Machine (SVMrbf\text{SVM}_{rbf}), a 3-layer Sparse Auto-Encoder (SAE-3), a 3-layer Restricted Boltzmann Machine (RBM-3), and a 3-layer Denoising Auto-Encoder with binary masking noise (DAE-b-3). Test error rates with 95% confidence intervals demonstrate that a 2-layer stacked CAE equals or outperforms 3-layer stacked models across most benchmarks:

    Dataset SVMrbf\text{SVM}_{rbf} SAE-3 RBM-3 DAE-b-3 CAE-1 CAE-2
    basic 3.03±0.153.03 \pm 0.15 3.46±0.163.46 \pm 0.16 3.11±0.153.11 \pm 0.15 2.84±0.152.84 \pm 0.15 2.83±0.152.83 \pm 0.15 2.48±0.14\mathbf{2.48 \pm 0.14}
    rot 11.11±0.2811.11 \pm 0.28 10.30±0.2710.30 \pm 0.27 10.30±0.2710.30 \pm 0.27 9.53±0.26\mathbf{9.53 \pm 0.26} 11.59±0.2811.59 \pm 0.28 9.66±0.26\mathbf{9.66 \pm 0.26}
    bg-rand 14.58±0.3114.58 \pm 0.31 11.28±0.2811.28 \pm 0.28 6.73±0.22\mathbf{6.73 \pm 0.22} 10.30±0.2710.30 \pm 0.27 13.57±0.3013.57 \pm 0.30 10.90±0.2710.90 \pm 0.27
    bg-img 22.61±0.3822.61 \pm 0.38 23.00±0.3723.00 \pm 0.37 16.31±0.3216.31 \pm 0.32 16.68±0.3316.68 \pm 0.33 16.70±0.3316.70 \pm 0.33 15.50±0.32\mathbf{15.50 \pm 0.32}
    bg-img-rot 55.18±0.4455.18 \pm 0.44 51.93±0.4451.93 \pm 0.44 47.39±0.4447.39 \pm 0.44 43.76±0.43\mathbf{43.76 \pm 0.43} 48.10±0.4448.10 \pm 0.44 45.23±0.4445.23 \pm 0.44
    rect 2.15±0.132.15 \pm 0.13 2.41±0.132.41 \pm 0.13 2.60±0.142.60 \pm 0.14 1.99±0.121.99 \pm 0.12 1.48±0.101.48 \pm 0.10 1.21±0.10\mathbf{1.21 \pm 0.10}
    rect-img 24.04±0.3724.04 \pm 0.37 24.05±0.3724.05 \pm 0.37 22.50±0.3722.50 \pm 0.37 21.59±0.36\mathbf{21.59 \pm 0.36} 21.86±0.3621.86 \pm 0.36 21.54±0.36\mathbf{21.54 \pm 0.36}
  10. Knowl 10 — Impact of Layer Stacking on Feature Invariance and Contraction Profile

    empirical result

    When multiple Contractive Auto-Encoder layers are trained greedily and stacked sequentially (where the hidden representation h(l−1)h^{(l-1)} of layer l−1l-1 serves as input to layer ll), the composite encoder function f(l)∘⋯∘f(1)f^{(l)} \circ \dots \circ f^{(1)} exhibits an enhanced contractive profile:

    1. Sharper Invariance Plateau: Stacking depth increases the degree of contraction (lowering the contraction ratio) at short and medium distances rr around training samples, forming a broader and flatter plateau of local invariance.
    2. Delayed Saturation: The distance rr at which the contraction ratio begins to rise due to global saturation effects is shifted farther away from the training points as depth increases from 1 to 2 and 3 layers.

    This demonstrates that deeper contractive architectures extract higher-level features with broader spatial invariance along non-linear manifold directions.

Coverage note — None was omitted; the knowls cover the CAE objective formulation, the analytic sigmoid Frobenius penalty and complexity, geometric manifold interpretation, theoretical links to DAE, weight decay, and sparse coding, contraction ratio definitions, singular value spectrum analysis, and all empirical benchmark results on MNIST, CIFAR-bw, and the ICML 2007 variation datasets.

References

  1. 1.Baldi, P. and Hornik, K. (1989). Neural networks and principal component analysis: Learning from examples without local minima. Neural Networks, 2, 53–58.
  2. 2.Bengio, Y. (2009). Learning deep architectures for AI. Foundations and Trends in Machine Learning, 2(1), 1–127. Also published as a book. Now Publishers, 2009.
  3. 3.Bengio, Y., Lamblin, P., Popovici, D., and Larochelle, H. (2007). Greedy layer-wise training of deep networks. In B. Sch¨olkopf, J. Platt, and T. Hoffman, editors, Advances in Neural Information Processing Systems 19 (NIPS’06), pages 153–160. MIT Press.
  4. 4.Bishop, C. M. (1995). Training with noise is equivalent to Tikhonov regularization. Neural Computation, 7(1), 108–116.
  5. 5.Hinton, G. E., Osindero, S., and Teh, Y. (2006). A fast learning algorithm for deep belief nets. Neural Computation, 18, 1527–1554.
  6. 6.Japkowicz, N., Hanson, S. J., and Gluck, M. A. (2000). Nonlinear autoassociation is not equivalent to PCA. Neural Computation, 12(3), 531–545.
  7. 7.Jarrett, K., Kavukcuoglu, K., Ranzato, M., and LeCun, Y. (2009). What is the best multi-stage architecture for object recognition? In Proc. International Conference on Computer Vision (ICCV’09). IEEE.
  8. 8.Kavukcuoglu, K., Ranzato, M., Fergus, R., and LeCun, Y. (2009). Learning invariant features through topographic filter maps. In Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR’09). IEEE.
  9. 9.Krizhevsky, A. and Hinton, G. (2009). Learning multiple layers of features from tiny images. Technical report, University of Toronto.
  10. 10.Larochelle, H., Erhan, D., Courville, A., Bergstra, J., and Bengio, Y. (2007). An empirical evaluation of deep architectures on problems with many factors of variation. In Z. Ghahramani, editor, Proceedings of the Twenty-fourth International Conference on Machine Learning (ICML’07), pages 473–480. ACM.
  11. 11.Lee, H., Ekanadham, C., and Ng, A. (2008). Sparse deep belief net model for visual area V2. In J. Platt, D. Koller, Y. Singer, and S. Roweis, editors, Advances in Neural Information Processing Systems 20 (NIPS’07), pages 873–880. MIT Press, Cambridge, MA.
  12. 12.Lee, H., Grosse, R., Ranganath, R., and Ng, A. Y. (2009). Convolutional deep belief networks for scalable unsupervised learning of hierarchical representations. In L. Bottou and M. Littman, editors, Proceedings of the Twenty-sixth International Conference on Machine Learning (ICML’09). ACM, Montreal (Qc), Canada.
  13. 13.Olshausen, B. A. and Field, D. J. (1997). Sparse coding with an overcomplete basis set: a strategy employed by V1? Vision Research, 37, 3311–3325.
  14. 14.Rumelhart, D. E., Hinton, G. E., and Williams, R. J. (1986). Learning representations by backpropagating errors. Nature, 323, 533–536.
  15. 15.Simard, P., Victorri, B., LeCun, Y., and Denker, J. (1992). Tangent prop - A formalism for specifying selected invariances in an adaptive network. In J. M. S. Hanson and R. Lippmann, editors, Advances in Neural Information Processing Systems 4 (NIPS’91), pages 895–903, San Mateo, CA. Morgan Kaufmann.
  16. 16.Vincent, P., Larochelle, H., Lajoie, I., Bengio, Y., and Manzagol, P.-A. (2010). Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion. Journal of Machine Learning Research, 11(3371–3408).
  17. 17.Weston, J., Ratle, F., and Collobert, R. (2008). Deep learning via semi-supervised embedding. In W. W. Cohen, A. McCallum, and S. T. Roweis, editors, Proceedings of the Twenty-fifth International Conference on Machine Learning (ICML’08), pages 1168–1175, New York, NY, USA. ACM.

Citation

MLA
Rifai, S., et al. “Contractive Auto-Encoders: Explicit Invariance During Feature Extraction”. International Conference on Machine Learning, 2011, pp. 833–40, http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.298.1581.
APA
Rifai, S., Vincent, P., Muller, X., Glorot, X., & Bengio, Y. (2011). Contractive Auto-Encoders: Explicit Invariance During Feature Extraction. International Conference on Machine Learning, 833–840. http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.298.1581
Chicago
Rifai, S., P. Vincent, X. Muller, X. Glorot, and Y. Bengio. 2011. “Contractive Auto-Encoders: Explicit Invariance During Feature Extraction”. International Conference on Machine Learning, 833–40. http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.298.1581.
Harvard
Rifai, S. et al. (2011) “Contractive Auto-Encoders: Explicit Invariance During Feature Extraction”, International Conference on Machine Learning, pp. 833–840. Available at: http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.298.1581.
Vancouver
1. Rifai S, Vincent P, Muller X, Glorot X, Bengio Y (2011) Contractive Auto-Encoders: Explicit Invariance During Feature Extraction. International Conference on Machine Learning 833–840

BibTeX

@article{rifai2011contractive,
  title = {Contractive Auto-Encoders: Explicit Invariance During Feature Extraction},
  author = {Rifai, Salah and Vincent, Pascal and Muller, Xavier and Glorot, Xavier and Bengio, Yoshua},
  year = {2011},
  journal = {International Conference on Machine Learning},
  pages = {833-840},
  url = {http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.298.1581}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission