Identifying and attacking the saddle point problem in high-dimensional non-convex optimization

Yann DauphinRazvan PascanuCaglar GulcehreKyunghyun ChoSurya GanguliYoshua Bengio

article2014NeurIPS1,555 citations

Demonstrates that saddle points, rather than local minima, dominate high-dimensional non-convex optimization and introduces a saddle-free Newton method designed to rapidly escape them during neural network training.

Listen

Minimizing complex, non-convex error functions in high-dimensional spaces is a fundamental challenge across machine learning and data science. Practitioners have long assumed that the primary bottleneck in training large models is getting trapped in sub-optimal local minima. Consequently, standard optimization tools—such as gradient descent and classical Newton-type methods—were designed around this assumption. However, these techniques frequently stall during training, leading to excessive compute costs, prolonged development cycles, and underperforming models.

The article demonstrates that the primary obstacle in high-dimensional optimization is not poor local minima, but rather an overwhelming proliferation of saddle points surrounded by flat, high-error plateaus. It introduces and evaluates a novel optimization algorithm, the saddle-free Newton method, designed to rapidly escape these saddle points and significantly accelerate model convergence.

The authors combined theoretical principles from statistical physics and random matrix theory with empirical evaluations on neural networks. They mapped the landscapes of multi-layer networks on standard image datasets (MNIST and CIFAR-10) using Newton's method to locate critical points across various error levels. To make second-order curvature calculations practical for larger architectures, they implemented the saddle-free Newton method within a lower-dimensional subspace using Krylov techniques. The algorithm was then tested against stochastic gradient descent and classical damped Newton methods on feedforward networks, deep autoencoders, and recurrent neural networks.

The findings confirm that high-error critical points in high-dimensional spaces are almost exclusively saddle points rather than local minima, with the proportion of negative curvature directions increasing directly with error. Standard gradient descent slows down severely on the flat plateaus surrounding these points, while classical Newton methods are actively attracted to them by moving in the wrong direction along negative curvatures. In contrast, the proposed saddle-free Newton method rescales gradients by the absolute values of the curvature matrix, repelling the optimizer away from saddle points. In empirical benchmarks, this approach broke through stagnation plateaus where gradient descent stalled. On a standard seven-layer deep autoencoder benchmark, it achieved a new state-of-the-art mean-squared error of 0.57, outperforming the prior benchmark of 0.69 set by Hessian-Free optimization, and substantially reduced underfitting in recurrent neural networks.

These insights fundamentally change our understanding of non-convex optimization. Organizations investing heavily in deep learning infrastructure can improve training efficiency and model accuracy by adopting algorithms specifically engineered to navigate saddle points rather than avoid non-existent high-error local minima. While exact implementations require higher computational overhead per step, the substantial reduction in total training iterations and superior final performance can significantly reduce computing costs and timeline risks.

Engineering and research teams should evaluate saddle-free optimization principles, particularly when scaling deep feedforward or recurrent network architectures that suffer from training plateaus. The primary limitation of the exact saddle-free formulation is its computational demand in high dimensions, necessitating approximate methods like Krylov subspaces. Future work should focus on scaling these saddle-free mechanisms to larger systems and exploring scalable alternatives beyond Krylov subspaces.

arXiv: 1406.2572
Cover for Identifying and attacking the saddle point problem in high-dimensional non-convex optimization

Abstract

A central challenge to many fields of science and engineering involves minimizing non-convex error functions over continuous, high dimensional spaces. Gradient descent or quasi-Newton methods are almost ubiquitously used to perform such minimizations, and it is often thought that a main source of difficulty for these local methods to find the global minimum is the proliferation of local minima with much higher error than the global minimum. Here we argue, based on results from statistical physics, random matrix theory, neural network theory, and empirical evidence, that a deeper and more profound difficulty originates from the proliferation of saddle points, not local minima, especially in high dimensional problems of practical interest. Such saddle points are surrounded by high error plateaus that can dramatically slow down learning, and give the illusory impression of the existence of a local minimum. Motivated by these arguments, we propose a new approach to second-order optimization, the saddle-free Newton method, that can rapidly escape high dimensional saddle points, unlike gradient descent and quasi-Newton methods. We apply this algorithm to deep or recurrent neural network training, and provide numerical evidence for its superior optimization performance.

Table of Contents

  • 1 Introduction
  • 2 The prevalence of saddle points in high dimensions
  • 3 Experimental validation of the prevalence of saddle points
  • 4 Dynamics of optimization algorithms near saddle points
  • 5 Generalized trust region methods
  • 6 Attacking the saddle point problem
  • 7 Experimental validation of the saddle-free Newton method
  • 7.1 Feedforward Neural Networks
  • 7.1.1 Existence of Saddle Points in Neural Networks
  • 7.1.2 Effectiveness of saddle-free Newton Method in Deep Neural Networks
  • 7.2 Recurrent Neural Networks: Hard Optimization Problem
  • 8 Conclusion
  • References
  • A Description of the different types of saddle-points
  • B Reparametrization of the space around saddle-points
  • C Empirical exploration of properties of critical points
  • D Proof of Lemma
  • E Implementation details for approximate saddle-free Newton
  • F Experiments
  • F.1 Existence of Saddle Points in Neural Networks
  • F.2 Effectiveness of saddle-free Newton Method in Deep Neural Networks
  • F.3 Recurrent Neural Networks: Hard Optimization Problem

Knowls

  1. Knowl 1 — Saddle-Free Newton Optimization Method

    model/method

    The saddle-free Newton (SFN) method is a second-order optimization algorithm designed to escape saddle points rapidly in non-convex optimization problems. For a scalar objective function f(θ)f(\theta) with continuous parameters θ∈RN\theta \in \mathbb{R}^N, gradient ∇f(θ)∈RN\nabla f(\theta) \in \mathbb{R}^N, and symmetric Hessian matrix H(θ)=∇2f(θ)∈RN×NH(\theta) = \nabla^2 f(\theta) \in \mathbb{R}^{N \times N} with spectral decomposition H=∑i=1Nλieiei⊤H = \sum_{i=1}^N \lambda_i e_i e_i^\top, the saddle-free Newton step is defined as:

    Δθ=−∇f(θ)∣H(θ)∣−1\Delta \theta = -\nabla f(\theta) |H(\theta)|^{-1}

    where ∣H(θ)∣|H(\theta)| is the absolute Hessian matrix obtained by taking the absolute value of each eigenvalue:

    ∣H(θ)∣=∑i=1N∣λi∣eiei⊤|H(\theta)| = \sum_{i=1}^N |\lambda_i| e_i e_i^\top

    In terms of the projection of the gradient onto each eigenvector eie_i, the step along direction eie_i is −1∣λi∣(ei⊤∇f(θ))ei-\frac{1}{|\lambda_i|} (e_i^\top \nabla f(\theta)) e_i.

    This update rule matches the classical Newton step along directions of positive curvature (where λi>0\lambda_i > 0), while along directions of negative curvature (where λi<0\lambda_i < 0), it preserves the descent direction of gradient descent by moving away from the critical point rather than towards it, converting saddle points from attractors into repellers while maintaining inverse-curvature step-size scaling.

  2. Knowl 2 — Generalized Trust Region Derivation of the Saddle-Free Newton Step

    theoretical result

    The saddle-free Newton step is derived from a generalized trust region optimization framework where the objective is approximated by a first-order Taylor expansion and the trust region constraint bounds the discrepancy between first-order and second-order Taylor expansions.

    The constrained optimization problem is defined as:

    Δθ=arg⁡min⁡Δθ(f(θ)+∇f(θ)⊤Δθ)subject tod(θ,θ+Δθ)≤Δ\Delta \theta = \arg \min_{\Delta \theta} \left( f(\theta) + \nabla f(\theta)^\top \Delta \theta \right) \quad \text{subject to} \quad d(\theta, \theta + \Delta \theta) \le \Delta

    The discrepancy metric is given by:

    d(θ,θ+Δθ)=∣f(θ)+∇f(θ)⊤Δθ+12Δθ⊤HΔθ−f(θ)−∇f(θ)⊤Δθ∣=12∣Δθ⊤HΔθ∣≤Δd(\theta, \theta + \Delta \theta) = \left| f(\theta) + \nabla f(\theta)^\top \Delta \theta + \frac{1}{2} \Delta \theta^\top H \Delta \theta - f(\theta) - \nabla f(\theta)^\top \Delta \theta \right| = \frac{1}{2} |\Delta \theta^\top H \Delta \theta| \le \Delta

    Using the upper bound ∣Δθ⊤HΔθ∣≤Δθ⊤∣H∣Δθ|\Delta \theta^\top H \Delta \theta| \le \Delta \theta^\top |H| \Delta \theta, the trust region problem is relaxed to:

    Δθ=arg⁡min⁡Δθ(f(θ)+∇f(θ)⊤Δθ)subject toΔθ⊤∣H∣Δθ≤Δ\Delta \theta = \arg \min_{\Delta \theta} \left( f(\theta) + \nabla f(\theta)^\top \Delta \theta \right) \quad \text{subject to} \quad \Delta \theta^\top |H| \Delta \theta \le \Delta

    Because the linear objective has its minimum at infinity, the optimal solution lies on the boundary (equality constraint). Solving via the method of Lagrange multipliers yields:

    Δθ=−∇f(θ)∣H∣−1\Delta \theta = -\nabla f(\theta) |H|^{-1}

    where the Lagrange multiplier scalar is absorbed into the global learning rate.

  3. Knowl 3 — Upper Bound on Indefinite Quadratic Forms via Matrix Absolute Value

    theoretical result

    Let A∈Rn×nA \in \mathbb{R}^{n \times n} be a nonsingular square symmetric matrix with eigendecomposition A=∑i=1nλieiei⊤A = \sum_{i=1}^n \lambda_i e_i e_i^\top, where λi∈R\lambda_i \in \mathbb{R} are eigenvalues and ei∈Rne_i \in \mathbb{R}^n are orthonormal eigenvectors. Let ∣A∣|A| denote the matrix obtained by replacing each eigenvalue with its absolute value:

    ∣A∣=∑i=1n∣λi∣eiei⊤|A| = \sum_{i=1}^n |\lambda_i| e_i e_i^\top

    For any vector x∈Rnx \in \mathbb{R}^n, the quadratic form satisfies the upper bound:

    ∣x⊤Ax∣≤x⊤∣A∣x|x^\top A x| \le x^\top |A| x

  4. Knowl 4 — Local Dynamics of Optimization Algorithms near Non-Degenerate Saddle Points

    theoretical result

    Near a non-degenerate critical point θ∗\theta^* where the gradient vanishes and the Hessian HH has non-zero eigenvalues λ1,…,λN\lambda_1, \dots, \lambda_N with orthonormal eigenvectors e1,…,eNe_1, \dots, e_N, the objective function f(θ)f(\theta) can be locally re-parameterized via Morse's lemma as:

    f(θ∗+Δθ)=f(θ∗)+12∑i=1NλiΔvi2f(\theta^* + \Delta \theta) = f(\theta^*) + \frac{1}{2} \sum_{i=1}^N \lambda_i \Delta v_i^2

    where Δvi=ei⊤Δθ\Delta v_i = e_i^\top \Delta \theta represents the displacement along eigenvector eie_i. Different optimization algorithms behave along each eigen-direction eie_i as follows:

    • Gradient Descent: Takes step −λiΔvi-\lambda_i \Delta v_i. It moves in the correct direction (toward θ∗\theta^* if λi>0\lambda_i > 0, away from θ∗\theta^* if λi<0\lambda_i < 0), but step sizes become vanishingly small along directions of small absolute curvature ∣λi∣≈0|\lambda_i| \approx 0.
    • Newton's Method: Takes step −Δvi-\Delta v_i. When λi<0\lambda_i < 0, the step is directed toward θ∗\theta^* (opposite to the descent direction), turning the saddle point into an attractor of the dynamics.
    • Damped Newton's Method: Adds damping α>0\alpha > 0 to the diagonal, taking step −λiλi+αΔvi-\frac{\lambda_i}{\lambda_i + \alpha} \Delta v_i. To guarantee descent along all negative directions, one must set α>∣λmin⁡∣\alpha > |\lambda_{\min}|, which severely reduces the effective step size −λiλi+α-\frac{\lambda_i}{\lambda_i + \alpha} along directions with small positive eigenvalues.
    • Saddle-Free Newton Method: Rescales by the inverse absolute eigenvalue, yielding the step −sgn⁡(λi)Δvi-\operatorname{sgn}(\lambda_i) \Delta v_i. This ensures fast escape along negative curvature directions without damping positive curvature directions.
  5. Knowl 5 — Approximate Saddle-Free Newton via Krylov Subspace Projection

    algorithm

    To apply the saddle-free Newton method in high dimensions where computing and inverting the full N×NN \times N Hessian is intractable, the objective f(θ)f(\theta) is optimized within a kk-dimensional Krylov subspace (k≪Nk \ll N) spanned by orthonormal vectors V=[v1,…,vk]∈RN×kV = [v_1, \dots, v_k] \in \mathbb{R}^{N \times k} generated via a modified Lanczos iteration that incorporates the gradient g=−∇f(θ)g = -\nabla f(\theta) and the previous update direction Δθ\Delta \theta.

    Input: Objective function f(θ)f(\theta), current parameter θ\theta, outer iterations MM, inner iterations mm, subspace dimension kk
    Output: Updated parameter θ\theta
    for i=1i = 1 to MM do
        Compute V∈RN×kV \in \mathbb{R}^{N \times k} via Lanczos iteration on ∇2f(θ)\nabla^2 f(\theta) using g=−∇f(θ)g = -\nabla f(\theta) and previous step Δθ\Delta \theta
        Define subspace function f^(α)=f(θ+Vα)\hat{f}(\alpha) = f(\theta + V \alpha) where α∈Rk\alpha \in \mathbb{R}^k
        Compute subspace Hessian H^=∇2f^(α)=V⊤∇2f(θ)V\hat{H} = \nabla^2 \hat{f}(\alpha) = V^\top \nabla^2 f(\theta) V
        Compute absolute subspace Hessian ∣H^∣=∑j=1k∣λj∣ujuj⊤|\hat{H}| = \sum_{j=1}^k |\lambda_j| u_j u_j^\top via eigendecomposition H^=∑j=1kλjujuj⊤\hat{H} = \sum_{j=1}^k \lambda_j u_j u_j^\top
        for j=1j = 1 to mm do
            Compute subspace gradient gα=−∇f^(α)=−V⊤∇f(θ)g_\alpha = -\nabla \hat{f}(\alpha) = -V^\top \nabla f(\theta)
            Find damping parameter λ=arg⁡min⁡λf^(gα(∣H^∣+λI)−1)\lambda = \arg \min_\lambda \hat{f}\left(g_\alpha (|\hat{H}| + \lambda I)^{-1}\right)
            Update parameter θ←θ+gα(∣H^∣+λI)−1V⊤\theta \leftarrow \theta + g_\alpha (|\hat{H}| + \lambda I)^{-1} V^\top
        end for
    end for
    return θ\theta

    Hessian-vector products Vj∇2f(θ)V_j \nabla^2 f(\theta) during the Lanczos process are computed via fast Pearlmutter L\mathcal{L}-operator passes without forming the full Hessian matrix.

  6. Knowl 6 — Empirical Correlation Between Critical Point Index and Error in Neural Networks

    empirical result

    Empirical analysis of the loss surfaces of multi-layer perceptrons trained on downsampled (10×1010 \times 10) MNIST and CIFAR-10 datasets reveals that critical points identified via Newton's method concentrate along a monotonically increasing curve in the error-index plane (ϵ,α)(\epsilon, \alpha), where ϵ\epsilon is the training error and α\alpha is the index (fraction of negative eigenvalues of the Hessian matrix at the critical point).

    Key observations include:

    • Critical points with high training error possess a high index α\alpha (ranging up to ≈0.25\approx 0.25), confirming that high-error stationary points are overwhelmingly saddle points rather than local minima.
    • Local minima (index α=0\alpha = 0) occur exclusively at low error levels close to the global minimum.
    • The Hessian eigenvalue distribution shifts rightward as error decreases, with a prominent mode centered at zero across all critical points, confirming the ubiquitous presence of small-curvature plateaus surrounding saddle points in neural network cost surfaces.
  7. Knowl 7 — Saddle-Free Newton Optimization of Deep Autoencoders on MNIST

    empirical result

    The saddle-free Newton (SFN) algorithm was evaluated on a 7-hidden-layer deep autoencoder benchmark on full-scale MNIST using Krylov subspace descent with k=500k = 500 subspace vectors.

    When initialized after stochastic gradient descent (SGD) stalled at a mean squared error (MSE) plateau of 1.0, SFN escaped the plateau rapidly, producing an accelerated rightward shift in the Hessian eigenvalue spectrum. SFN achieved a final test MSE of 0.57, outperforming the previous state-of-the-art benchmark of 0.69 MSE achieved by Hessian-Free optimization on the identical architecture.

  8. Knowl 8 — Saddle-Free Newton Optimization of Recurrent Neural Networks

    empirical result

    The saddle-free Newton (SFN) method was applied to train a recurrent neural network with 120 hidden units on character-level language modeling on the Penn Treebank corpus.

    After standard SGD stalled due to optimization plateaus, continuation with SFN resulted in a substantial reduction in training error. In contrast, truncated Newton methods with diagonal damping failed to make progress from the stalled SGD point. Spectral analysis of the Krylov subspace Hessian confirmed that the solutions reached by SFN had significantly fewer negative eigenvalues compared to the stalled states of SGD.

  9. Knowl 9 — Dimensionality-Dependent Performance Scaling of Saddle-Free Newton

    empirical result

    Evaluation of minibatch stochastic gradient descent (MSGD), damped Newton, and exact saddle-free Newton (SFN) on single-layer MLPs of varying widths (5, 25, and 50 hidden units) trained on 10×1010 \times 10 downsampled MNIST and CIFAR-10 demonstrates that optimization difficulty increases with parameter dimensionality:

    • At 5 hidden units, MSGD, damped Newton, and SFN attain comparable final training error.
    • As network width increases to 25 and 50 hidden units (expanding parameter space dimensionality), SFN outperforms both MSGD and damped Newton by substantial margins.
    • MSGD and damped Newton get trapped near high-error saddle points early in training (e.g., around epoch 10 on MNIST), whereas SFN escapes these plateaus and drives training error significantly lower.

Coverage note — Omitted prior literature reviews on replica theory in Gaussian random fields, theoretical properties of linear MLPs, and intermediate proof manipulations of Lemma 1 to concentrate strictly on the paper's novel contributions, algorithms, theoretical derivations, and empirical validations.

References

  1. 1.Baldi, P. and Hornik, K. (1989). Neural networks and principal component analysis: Learning from examples without local minima. Neural Networks, 2(1), 53–58.
  2. 2.Bastien, F., Lamblin, P., Pascanu, R., Bergstra, J., Goodfellow, I. J., Bergeron, A., Bouchard, N., and Bengio, Y. (2012). Theano: new features and speed improvements.
  3. 3.Bengio, Y., Simard, P., and Frasconi, P. (1994). Learning long-term dependencies with gradient descent is difficult. 5(2), 157–166. Special Issue on Recurrent Neural Networks, March 94.
  4. 4.Bergstra, J. and Bengio, Y. (2012). Random search for hyper-parameter optimization. Journal of Machine Learning Research, 13, 281–305.
  5. 5.Bergstra, J., Breuleux, O., Bastien, F., Lamblin, P., Pascanu, R., Desjardins, G., Turian, J., Warde-Farley, D., and Bengio, Y. (2010). Theano: a CPU and GPU math expression compiler. In Proceedings of the Python for Scientific Computing Conference (SciPy).
  6. 6.Bray, A. J. and Dean, D. S. (2007). Statistics of critical points of gaussian fields on large-dimensional spaces. Physics Review Letter, 98, 150201.
  7. 7.Callahan, J. (2010). Advanced Calculus: A Geometric View. Undergraduate Texts in Mathematics. Springer.
  8. 8.Fyodorov, Y. V. and Williams, I. (2007). Replica symmetry breaking condition exposed by random matrix calculation of landscape complexity. Journal of Statistical Physics, 129(5-6), 1081–1116.
  9. 9.Inoue, M., Park, H., and Okada, M. (2003). On-line learning theory of soft committee machines with correlated hidden units steepest gradient descent and natural gradient descent. Journal of the Physical Society of Japan, 72(4), 805–810.
  10. 10.Le Roux, N., Manzagol, P.-A., and Bengio, Y. (2007). Topmoumoute online natural gradient algorithm. Advances in Neural Information Processing Systems.
  11. 11.Martens, J. (2010). Deep learning via hessian-free optimization. In International Conference in Machine Learning, pages 735–742.
  12. 12.Mizutani, E. and Dreyfus, S. (2010). An analysis on negative curvature induced by singularity in multi-layer neural-network learning. In Advances in Neural Information Processing Systems, pages 1669–1677.
  13. 13.Murray, W. (2010). Newton-type methods. Technical report, Department of Management Science and Engineering, Stanford University.
  14. 14.Nocedal, J. and Wright, S. (2006). Numerical Optimization. Springer.
  15. 15.Parisi, G. (2007). Mean field theory of spin glasses: statistics and dynamics. Technical Report Arxiv 0706.0094.
  16. 16.Pascanu, R. and Bengio, Y. (2014). Revisiting natural gradient for deep networks. In International Conference on Learning Representations.
  17. 17.Pascanu, R., Mikolov, T., and Bengio, Y. (2013). On the difficulty of training recurrent neural networks. In ICML’2013.
  18. 18.Pascanu, R., Dauphin, Y., Ganguli, S., and Bengio, Y. (2014). On the saddle point problem for non-convex optimization. Technical Report Arxiv 1405.4604.
  19. 19.Pearlmutter, B. A. (1994). Fast exact multiplication by the hessian. Neural Computation, 6, 147–160.
  20. 20.Rasmussen, C. E. and Williams, C. K. I. (2005). Gaussian Processes for Machine Learning (Adaptive Computation and Machine Learning). The MIT Press.
  21. 21.Rattray, M., Saad, D., and Amari, S. I. (1998). Natural Gradient Descent for On-Line Learning. Physical Review Letters, 81(24), 5461–5464.
  22. 22.Saad, D. and Solla, S. A. (1995). On-line learning in soft committee machines. Physical Review E, 52, 4225–4243.
  23. 23.Saxe, A., McClelland, J., and Ganguli, S. (2013). Learning hierarchical category structure in deep neural networks. Proceedings of the 35th annual meeting of the Cognitive Science Society, pages 1271–1276.
  24. 24.Saxe, A., McClelland, J., and Ganguli, S. (2014). Exact solutions to the nonlinear dynamics of learning in deep linear neural network. In International Conference on Learning Representations.
  25. 25.Sohl-Dickstein, J., Poole, B., and Ganguli, S. (2014). Fast large-scale optimization by unifying stochastic gradient and quasi-newton methods. In ICML’2014.
  26. 26.Sutskever, I., Martens, J., Dahl, G. E., and Hinton, G. E. (2013). On the importance of initialization and momentum in deep learning. In S. Dasgupta and D. Mcallester, editors, Proceedings of the 30th International Conference on Machine Learning (ICML-13), volume 28, pages 1139–1147. JMLR Workshop and Conference Proceedings.
  27. 27.Vinyals, O. and Povey, D. (2012). Krylov Subspace Descent for Deep Learning. In AISTATS.
  28. 28.Wigner, E. P. (1958). On the distribution of the roots of certain symmetric matrices. The Annals of Mathematics, 67(2), 325–327.

Citation

MLA
Dauphin, Y., et al. “Identifying and Attacking the Saddle Point Problem in High-dimensional Non-convex Optimization”. arXiv, 2014, http://arxiv.org/abs/1406.2572v1.
APA
Dauphin, Y., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., & Bengio, Y. (2014). Identifying and attacking the saddle point problem in high-dimensional non-convex optimization. arXiv. http://arxiv.org/abs/1406.2572v1
Chicago
Dauphin, Y., R. Pascanu, C. Gulcehre, K. Cho, S. Ganguli, and Y. Bengio. 2014. “Identifying and Attacking the Saddle Point Problem in High-dimensional Non-convex Optimization”. arXiv. http://arxiv.org/abs/1406.2572v1.
Harvard
Dauphin, Y. et al. (2014) “Identifying and attacking the saddle point problem in high-dimensional non-convex optimization”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1406.2572v1.
Vancouver
1. Dauphin Y, Pascanu R, Gulcehre C, Cho K, Ganguli S, Bengio Y (2014) Identifying and attacking the saddle point problem in high-dimensional non-convex optimization. arXiv

BibTeX

@article{dauphin2014identifying,
  title = {Identifying and attacking the saddle point problem in high-dimensional non-convex optimization},
  author = {Dauphin, Yann and Pascanu, Razvan and Gulcehre, Caglar and Cho, Kyunghyun and Ganguli, Surya and Bengio, Yoshua},
  year = {2014},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1406.2572v1},
  eprint = {1406.2572}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission