Deeply-Supervised Nets

Chen-Yu LeeSaining XiePatrick GallagherZhengyou ZhangZhuowen Tu

article2014AISTATS2,536 citationsTest-of-Time Award

Proposes a deep supervision architecture that injects companion loss functions directly into intermediate hidden layers, mitigating vanishing gradients and forcing early layers to learn highly discriminative features for image classification.

Listen

The article addresses challenges in training deep convolutional neural networks, including limited transparency and discriminativeness of features in hidden layers, exploding or vanishing gradients that slow convergence, and heavy reliance on large training datasets. These issues reduce reliability and effectiveness in image classification tasks despite the overall promise of deep learning.

The work evaluates a new formulation called deeply-supervised nets that adds direct supervision to hidden layers to improve feature quality and training efficiency while maintaining focus on final classification accuracy.

The approach augments standard CNN architectures by attaching companion classifiers, such as SVM or softmax, to each hidden layer. These produce additional loss terms in the overall objective, optimized jointly via stochastic gradient descent on benchmark datasets including MNIST, CIFAR-10, CIFAR-100, and SVHN, with model complexity matched to prior work and no data augmentation in the primary reported results.

Key findings show state-of-the-art classification errors of 0.39 percent on MNIST, 9.78 percent on CIFAR-10 without augmentation, 34.57 percent on CIFAR-100, and 1.92 percent on SVHN. DSN variants outperform corresponding CNN baselines, with gains reaching 26 percent relative improvement on MNIST when training data is limited to 500 samples. Features learned in early layers appear more intuitive, training converges faster, and gradient variance in the first layer increases by a factor of about 4.5. Generalization improves even when training error reaches near zero for both methods.

These results indicate that direct hidden-layer supervision acts as effective regularization, reduces dependence on massive datasets, and mitigates gradient problems without compromising output-layer performance. The gains matter for practical deployment where data is scarce, training time is costly, or robustness to hyperparameter choices is needed.

The formulation is compatible with existing techniques such as dropout and averaging, suggesting further error reductions are possible with engineering refinements. Additional validation on larger or more diverse tasks would strengthen confidence before broad adoption.

The main limitations include a convergence analysis that assumes local strong convexity around the optimum and experiments conducted without model averaging or extensive data augmentation in the core comparisons. Results are consistent across four standard benchmarks, supporting moderate confidence in the reported improvements for similar image-classification settings.

arXiv: 1409.5185s9xie/DSN
  • Paper: Understanding the difficulty of training deep feedforward neural networks, Xavier Glorot et al. (2010). This paper analyzes the vanishing gradient problems and training difficulties in deep feedforward architectures that deeply-supervised nets directly aim to resolve through intermediate companion losses.
  • Paper: Network In Network, Min Lin et al. (2014). This work introduces the Network In Network architecture and global average pooling, establishing key baseline architectures and micro-network components compared and built upon by deeply-supervised nets.
  • Paper: Visualizing and Understanding Convolutional Networks, Matthew D. Zeiler et al. (2014). This paper establishes deconvolutional feature visualization techniques to interpret intermediate representations in convolutional networks, motivating deeply-supervised nets to improve early-layer feature transparency and discriminativeness.
  • Paper: ImageNet Classification with Deep Convolutional Neural Networks, Alex Krizhevsky et al. (2012). This foundational work establishes the core deep convolutional neural network architecture and optimization practices on which deeply-supervised learning schemes are developed.
  • Paper: Greedy Layer-Wise Training of Deep Networks, Yoshua Bengio et al. (2007). This work introduces layer-wise training strategies for deep architectures, providing foundational concepts for providing direct training signals to hidden layers.
  • Paper: Improving neural networks by preventing co-adaptation of feature detectors, Geoffrey E. Hinton et al. (2012). This paper introduces dropout regularization, a standard baseline and complementary regularizer used alongside companion layer supervision in deeply-supervised networks.
  • Paper: FitNets: Hints for Thin Deep Nets, Adriana Romero et al. (2015). This work extends intermediate layer supervision by using teacher-student hints at hidden layers to train deeper and thinner neural networks.
  • Paper: Deep Residual Learning for Image Recognition, Kaiming He et al. (2016). This paper introduces residual learning with shortcut connections as an architectural solution to the degradation and gradient vanishing problems addressed by deeply-supervised nets.
  • Paper: Densely Connected Convolutional Networks, Gao Huang et al. (2017). This architecture advances direct feature reuse and gradient propagation throughout all hidden layers via dense feed-forward connections.
  • Paper: YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information, Chien-Yao Wang et al. (2024). This work builds upon auxiliary supervision and gradient flow concepts by introducing programmable gradient information via auxiliary branches during training.
  • Paper: Deep Networks with Stochastic Depth, Gao Huang et al. (2016). This paper explores stochastic depth as an alternative regularization and gradient-propagation technique for training very deep networks.
  • Paper: Visualizing the Loss Landscape of Neural Nets, Hao Li et al. (2017). This study analyzes and visualizes how network depth and structural modifications alter the loss landscape geometry and trainability of deep models.
Cover for Deeply-Supervised Nets

Abstract

Our proposed deeply-supervised nets (DSN) method simultaneously minimizes classification error while making the learning process of hidden layers direct and transparent. We make an attempt to boost the classification performance by studying a new formulation in deep networks. Three aspects in convolutional neural networks (CNN) style architectures are being looked at: (1) transparency of the intermediate layers to the overall classification; (2) discriminativeness and robustness of learned features, especially in the early layers; (3) effectiveness in training due to the presence of the exploding and vanishing gradients. We introduce "companion objective" to the individual hidden layers, in addition to the overall objective at the output layer (a different strategy to layer-wise pre-training). We extend techniques from stochastic gradient methods to analyze our algorithm. The advantage of our method is evident and our experimental result on benchmark datasets shows significant performance gain over existing methods (e.g. all state-of-the-art results on MNIST, CIFAR-10, CIFAR-100, and SVHN).

Table of Contents

  • 1 Introduction
  • 2 Deeply-Supervised Nets
  • 2.1 Motivation
  • 2.2 Formulation
  • 2.3 Stochastic Gradient Descent View
  • 3 Experiments
  • 3.1 MNIST
  • 3.2 CIFAR-10 and CIFAR-100
  • 3.3 Street View House Numbers
  • 4 Conclusions
  • 5 Acknowledgments
  • References

Knowls

  1. Knowl 1 — Deeply-Supervised Nets Formulation and Objective Function

    model/method

    In conventional deep convolutional neural networks, parameters W=(W(1),…,W(M))W = (W^{(1)}, \dots, W^{(M)}) across MM layers are guided solely by error backpropagated from the output layer. Deeply-Supervised Nets (DSN) introduce companion local classifiers w(m)w^{(m)} directly at each intermediate hidden layer m∈{1,…,M−1}m \in \{1, \dots, M-1\} alongside the standard output classifier w(out)w^{(\text{out})} at layer MM.

    Given an input sample X∈RnX \in \mathbb{R}^n with ground truth class label y∈{1,…,K}y \in \{1, \dots, K\}, layer representations are formed recursively by: Z(m)=f(Q(m)),Z(0)≡X,Q(m)=W(m)∗Z(m−1)Z^{(m)} = f(Q^{(m)}), \quad Z^{(0)} \equiv X, \quad Q^{(m)} = W^{(m)} * Z^{(m-1)} where W(m)W^{(m)} denotes the filters/weights at layer mm, Q(m)Q^{(m)} is the convolved feature response, f(⋅)f(\cdot) is a pooling and activation function, and Z(m)Z^{(m)} is the output feature map of layer mm.

    The overall combined objective function F(W)F(W) is defined as: F(W)≡P(W)+Q(W)=∥w(out)∥2+L(W,w(out))+∑m=1M−1αm[∥w(m)∥2+ℓ(W,w(m))−γ]+F(W) \equiv P(W) + Q(W) = \|w^{(\text{out})}\|^2 + \mathcal{L}(W, w^{(\text{out})}) + \sum_{m=1}^{M-1} \alpha_m \left[ \|w^{(m)}\|^2 + \ell(W, w^{(m)}) - \gamma \right]_+ where:

    • P(W)=∥w(out)∥2+L(W,w(out))P(W) = \|w^{(\text{out})}\|^2 + \mathcal{L}(W, w^{(\text{out})}) denotes the overall output layer objective with classifier weights w(out)w^{(\text{out})}.
    • Q(W)=∑m=1M−1αm[∥w(m)∥2+ℓ(W,w(m))−γ]+Q(W) = \sum_{m=1}^{M-1} \alpha_m [\|w^{(m)}\|^2 + \ell(W, w^{(m)}) - \gamma]_+ represents the companion objective across hidden layers with classifier weights w(m)w^{(m)}.
    • In an L2-SVM formulation with feature representation ϕ(⋅,⋅)\phi(\cdot, \cdot), the output loss L\mathcal{L} and companion loss ℓ\ell are squared hinge losses: L(W,w(out))=∑yk≠y[1−⟨w(out),ϕ(Z(M),y)−ϕ(Z(M),yk)⟩]+2\mathcal{L}(W, w^{(\text{out})}) = \sum_{y_k \neq y} \left[ 1 - \langle w^{(\text{out})}, \phi(Z^{(M)}, y) - \phi(Z^{(M)}, y_k) \rangle \right]_+^2 ℓ(W,w(m))=∑yk≠y[1−⟨w(m),ϕ(Z(m),y)−ϕ(Z(m),yk)⟩]+2\ell(W, w^{(m)}) = \sum_{y_k \neq y} \left[ 1 - \langle w^{(m)}, \phi(Z^{(m)}, y) - \phi(Z^{(m)}, y_k) \rangle \right]_+^2
    • [z]+=max⁡(0,z)[z]_+ = \max(0, z) is the hinge function. The threshold hyperparameter γ\gamma ensures that once the margin and prediction error at a hidden layer reach or drop below γ\gamma, the companion objective for that layer vanishes, acting as a feature regularization term without distorting final convergence.
    • αm\alpha_m is a weighting coefficient balancing companion supervision against the output loss, which can optionally decay across training epochs t∈{1,…,N}t \in \{1, \dots, N\} according to the schedule αm×0.1×(1−t/N)→αm\alpha_m \times 0.1 \times (1 - t/N) \to \alpha_m.
  2. Knowl 2 — Parameter Gradients for Deeply-Supervised Nets

    equation

    For a Deeply-Supervised Net using L2-SVM companion and output classifiers, the analytic gradients of the objective function F(W)F(W) with respect to the classifier weights are given by:

    Gradient with respect to the output classifier weights w(out)w^{(\text{out})}: ∂F∂w(out)=2w(out)−2∑yk≠y[ϕ(Z(M),y)−ϕ(Z(M),yk)][1−⟨w(out),ϕ(Z(M),y)−ϕ(Z(M),yk)⟩]+\frac{\partial F}{\partial w^{(\text{out})}} = 2 w^{(\text{out})} - 2 \sum_{y_k \neq y} [\phi(Z^{(M)}, y) - \phi(Z^{(M)}, y_k)] \left[ 1 - \langle w^{(\text{out})}, \phi(Z^{(M)}, y) - \phi(Z^{(M)}, y_k) \rangle \right]_+

    Gradient with respect to the companion classifier weights w(m)w^{(m)} at layer m∈{1,…,M−1}m \in \{1, \dots, M-1\}: ∂F∂w(m)={αm(2w(m)−2∑yk≠y[ϕ(Z(m),y)−ϕ(Z(m),yk)][1−⟨w(m),ϕ(Z(m),y)−ϕ(Z(m),yk)⟩]+),if ∥w(m)∥2+ℓ(W,w(m))>γ0,if ∥w(m)∥2+ℓ(W,w(m))≤γ\frac{\partial F}{\partial w^{(m)}} = \begin{cases} \alpha_m \left( 2 w^{(m)} - 2 \sum_{y_k \neq y} [\phi(Z^{(m)}, y) - \phi(Z^{(m)}, y_k)] \left[ 1 - \langle w^{(m)}, \phi(Z^{(m)}, y) - \phi(Z^{(m)}, y_k) \rangle \right]_+ \right), & \text{if } \|w^{(m)}\|^2 + \ell(W, w^{(m)}) > \gamma \\ 0, & \text{if } \|w^{(m)}\|^2 + \ell(W, w^{(m)}) \le \gamma \end{cases}

    The total gradient with respect to the convolutional weights WW is the sum of the conventional gradient backpropagated from the output classifier and the gradients propagated directly from each active companion objective whose total layer error exceeds γ\gamma.

  3. Knowl 3 — Layerwise Feasibility Propagation in Deep Supervision

    theoretical result

    In a Deeply-Supervised Net with MM layers and threshold γ\gamma, achieving a feasible solution for the companion loss at layer mm implies the existence of a feasible solution for all subsequent layers m′>mm' > m.

    Formally, for any layers m,m′∈{1,…,M−1}m, m' \in \{1, \dots, M-1\} with m′>mm' > m, if there exist learned weights (W^(1),…,W^(m))(\hat{W}^{(1)}, \dots, \hat{W}^{(m)}) and classifier w(m)w^{(m)} satisfying: ∥w(m)∥2+ℓ((W^(1),…,W^(m)),w(m))≤γ\|w^{(m)}\|^2 + \ell((\hat{W}^{(1)}, \dots, \hat{W}^{(m)}), w^{(m)}) \le \gamma then there exists an extension (W^(1),…,W^(m),…,W^(m′))(\hat{W}^{(1)}, \dots, \hat{W}^{(m)}, \dots, \hat{W}^{(m')}) and companion classifier w(m′)w^{(m')} such that: ∥w(m′)∥2+ℓ((W^(1),…,W^(m),…,W^(m′)),w(m′))≤γ\|w^{(m')}\|^2 + \ell((\hat{W}^{(1)}, \dots, \hat{W}^{(m)}, \dots, \hat{W}^{(m')}), w^{(m')}) \le \gamma

    This ensures that a feasible solution for the companion objective Q(W)Q(W) also yields a feasible solution for the output objective P(W)P(W), supporting the assumption that F(W)=P(W)+Q(W)F(W) = P(W) + Q(W) and P(W)P(W) share the same optimal parameter set W∗W^*.

  4. Knowl 4 — SGD Convergence Acceleration under Deep Supervision

    theoretical result

    Under the assumption of local strong convexity near the optimal solution W∗W^*, deeply-supervised nets provably accelerate the convergence rate of stochastic gradient descent (SGD) relative to standard top-layer-only supervision.

    Let the output loss P(W)P(W) be λ1\lambda_1-strongly convex and the companion objective Q(W)Q(W) be λ2\lambda_2-strongly convex near W∗W^*, making F(W)=P(W)+Q(W)F(W) = P(W) + Q(W) locally (λ1+λ2)(\lambda_1 + \lambda_2)-strongly convex. Assume subgradient second moments are bounded by E[∥g^P,t∥2]≤G2\mathbb{E}[\|\hat{g}_{P, t}\|^2] \le G^2 and E[∥g^Q,t∥2]≤G2\mathbb{E}[\|\hat{g}_{Q, t}\|^2] \le G^2, with initial distance ∥W1−W∗∥2≤D\|W_1 - W^*\|^2 \le D. Let WT(F)W_T^{(F)} and WT(P)W_T^{(P)} be the parameter estimates after TT iterations of SGD on F(W)F(W) and P(W)P(W), respectively.

    1. For step size ηt=1/(λt)=1/((λ1+λ2)t)\eta_t = 1/(\lambda t) = 1/((\lambda_1 + \lambda_2)t), the upper bound on expected parameter error yields the convergence speedup ratio: E[∥WT(P)−W∗∥2]E[∥WT(F)−W∗∥2]=Θ(1+λ22λ12)\frac{\mathbb{E}[\|W_T^{(P)} - W^*\|^2]}{\mathbb{E}[\|W_T^{(F)} - W^*\|^2]} = \Theta\left(1 + \frac{\lambda_2^2}{\lambda_1^2}\right)

    2. For step size ηt=1/t\eta_t = 1/t, the convergence speedup ratio is: E[∥WT(P)−W∗∥2]E[∥WT(F)−W∗∥2]=Θ(eλ2ln⁡T)=Θ(Tλ2)\frac{\mathbb{E}[\|W_T^{(P)} - W^*\|^2]}{\mathbb{E}[\|W_T^{(F)} - W^*\|^2]} = \Theta\left(e^{\lambda_2 \ln T}\right) = \Theta\left(T^{\lambda_2}\right)

    Since typically λ2≫λ1\lambda_2 \gg \lambda_1, the companion objective significantly accelerates the optimization convergence rate.

  5. Knowl 5 — DSN Network Architecture and Training Configuration

    experimental setup

    The DSN framework is implemented on top of the Caffe infrastructure with the following setup:

    • Network Backbone: Deep convolutional architecture utilizing Network in Network (NIN) style mlpconv layers and global average pooling, matching baseline model parameter complexity.
    • Regularization: Two dropout layers with a dropout rate of 0.50.5.
    • Classifier Variants: Intermediate and output classifiers are tested with both Softmax loss (DSN-Softmax) and linear Support Vector Machines using L2 squared hinge loss (DSN-SVM).
    • Optimization: Stochastic Gradient Descent (SGD) with a mini-batch size of 128 and a fixed momentum of 0.90.9.
    • Learning Rate Schedule: Initial learning rate and weight decay are tuned via validation sets. The learning rate is annealed by a factor of 20 according to validation performance.
    • Companion Error Propagation: Classification error signals from intermediate companion classifiers are directly backpropagated into their respective underlying convolutional layers.
  6. Knowl 6 — MNIST Classification Performance and Generalization

    data/table

    DSN was evaluated on the MNIST handwritten digit benchmark (60,000 training images, 10,000 testing images, 10 classes, 28×2828 \times 28 pixels) without data whitening or data augmentation.

    Method Error (%)
    CNN (Jarrett et al., 2009) 0.53
    Stochastic Pooling (Zeiler Fergus, 2013) 0.47
    Network in Network (Lin et al., 2014) 0.47
    Maxout Networks (Goodfellow et al., 2013) 0.45
    DSN (ours) 0.39

    Key observations:

    • DSN-SVM achieved an error rate of 0.39%0.39\% (DSN-Softmax achieved 0.51%0.51\%), improving over CNN-SVM (0.50%0.50\%) and CNN-Softmax (0.56%0.56\%).
    • Sample efficiency: When evaluated on reduced training subsets, DSN-SVM achieved a 26%26\% error reduction over CNN-Softmax at 500 training samples.
    • Generalization: While both standard CNN and DSN reached near-zero training error (0.03%0.03\%), DSN achieved lower test error (0.39%0.39\% vs 0.50%0.50\%), confirming the regularizing effect of intermediate supervision.
  7. Knowl 7 — CIFAR-10 and CIFAR-100 Classification Performance

    data/table

    DSN was evaluated on CIFAR-10 (10 classes, 5,000 training images/class) and CIFAR-100 (100 classes, 500 training images/class) using 32×3232 \times 32 images preprocessed with global contrast normalization. Data augmentation included 4-pixel zero padding, random corner cropping, and horizontal flipping. Testing was performed on single models using center crops without test-time model averaging.

    Method CIFAR-10 Error (%) CIFAR-100 Error (%)
    Without Data Augmentation
    Stochastic Pooling (Zeiler Fergus, 2013) 15.13 42.51
    Maxout Networks (Goodfellow et al., 2013) 11.68 38.57
    Network in Network (Lin et al., 2014) 10.41 35.68
    Tree based Priors (Srivastava Salakhutdinov, 2013) — 36.85
    DSN (ours) 9.78 34.57
    With Data Augmentation
    Maxout Networks (Goodfellow et al., 2013) 9.38 —
    DropConnect (Li et al., 2013) 9.32 —
    Network in Network (Lin et al., 2014) 8.81 —
    DSN (ours) 8.22 —

    DSN established new state-of-the-art results across both datasets, achieving 9.78%9.78\% (unaugmented) and 8.22%8.22\% (augmented) on CIFAR-10, and 34.57%34.57\% on CIFAR-100.

  8. Knowl 8 — Street View House Numbers Classification Performance

    data/table

    DSN was tested on the Street View House Numbers (SVHN) dataset (32×3232 \times 32 digit images). The training data included 598,388 images (combining 73,257 regular and 531,131 extra training images minus a validation split of 400 regular and 200 extra samples per class), with 26,032 test digits. Data was preprocessed with Local Contrast Normalization (LCN) without data augmentation during training or model voting during testing.

    Method Error (%)
    Stochastic Pooling (Zeiler Fergus, 2013) 2.80
    Maxout Networks (Goodfellow et al., 2013) 2.47
    Network in Network (Lin et al., 2014) 2.35
    DropConnect (Li et al., 2013) 1.94
    DSN (ours) 1.92

    DSN obtained an error rate of 1.92%1.92\% with a single model and no data augmentation, improving over competing methods including DropConnect (1.94%1.94\%, which used data augmentation and multi-model voting).

  9. Knowl 9 — Gradient Variance and Early Layer Representation Quality

    empirical result

    Companion objective supervision directly modifies gradient dynamics and feature representations in the bottom layers of the network:

    • Gradient Variance: In the first convolutional layer, DSN exhibits 4.554.55 times greater gradient variance than a standard CNN. This counteracts the vanishing gradient problem in early layers, facilitating faster training convergence and reducing sensitivity to hyperparameter initialization.
    • Representation Interpretability: Visualizing the top 30%30\% activations of first convolutional layer feature maps on CIFAR-10 shows that DSN learns cleaner, more intuitive, and more discriminative edge and pattern filters than a standard CNN trained with output supervision alone.

Coverage note — None. All main contributions—the mathematical formulation of DSN, gradient derivation, feasibility lemma, SGD convergence acceleration theorem, experimental protocol, benchmark classification tables (MNIST, CIFAR-10, CIFAR-100, SVHN), and empirical gradient/feature analyses—are fully represented.

References

  1. 1.Y. Bengio, P. Lamblin, D. Popovici, H. Larochelle, U. D. Montral, and M. Qubec. Greedy layer-wise training of deep networks. In NIPS, 2007.
  2. 2.J. Bergstra, O. Breuleux, F. Bastien, P. Lamblin, R. Pascanu, G. Desjardins, J. Turian, D. Warde-Farley, and Y. Bengio. Theano: a CPU and GPU math expression compiler. In Proceedings of the Python for Scientific Computing Conference (SciPy), June 2010.
  3. 3.L. Bottou. Online algorithms and stochastic approximations. Cambridge University Press, 1998.
  4. 4.G. E. Dahl, D. Yu, L. Deng, and A. Acero. Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition. IEEE Tran. on Audio, Speech, and Lang. Proc., 20(1):30–42, 2012.
  5. 5.J. Donahue, Y. Jia, O. Vinyals, J. Hoffman, N. Zhang, E. Tzeng, and T. Darrell. Decaf: A deep convolutional activation feature for generic visual recognition. In arXiv, 2013.
  6. 6.D. Eigen, J. Rolfe, R. Fergus, and Y. LeCun. Understanding deep architectures using a recursive convolutional network. In arXiv:1312.1847v2, 2014.
  7. 7.J. L. Elman. Distributed representations, simple recurrent networks, and grammatical. Machine Learning, 7:195–225, 1991.
  8. 8.X. Glorot and Y. Bengio. Understanding the difficulty of training deep feedforward neural networks. In AISTAT, 2010.
  9. 9.I. J. Goodfellow, D. Warde-Farley, M. Mirza, A. C. Courville, and Y. Bengio. Maxout networks. In ICML, 2013.
  10. 10.G. E. Hinton, S. Osindero, and Y. W. Teh. A fast learning algorithm for deep belief nets. Neural computation, 18:1527–1554, 2006.
  11. 11.G. E. Hinton, N. Srivastava, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov. Improving neural networks by preventing co-adaptation of feature detectors. In CoRR, abs/1207.0580, 2012.
  12. 12.F. J. Huang and Y. LeCun. Large-scale learning with svm and convolutional for generic object categorization. In CVPR, 2006.
  13. 13.K. Jarrett, K. Kavukcuoglu, M. Ranzato, and Y. LeCun. What is the best multi-stage architecture for object recognition? In ICCV, 2009.
  14. 14.Y. Jia. Caffe: An open source convolutional architecture for fast feature embedding. http://caffe.berkeleyvision.org/, 2013.
  15. 15.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In NIPS, 2012.
  16. 16.Q. Le, J. Ngiam, Z. Chen, D. Chia, P. W. Koh, and A. Ng. Tiled convolutional neural networks. In NIPS, 2010.
  17. 17.Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner. Gradient-based learning applied to document recognition. In Proceedings of the IEEE, 1998.
  18. 18.H. Lee, R. Grosse, R. Ranganath, and A. Y. Ng. Convolutional deep belief networks for scalable unsupervised learning of hierarchical representations. In ICML, 2009.
  19. 19.W. Li, M. Zeiler, S. Zhang, Y. LeCun, and R. Fergus. Regularization of neural networks using dropconnect. In ICML, 2013.
  20. 20.M. Lin, Q. Chen, and S. Yan. Network in network. In ICLR, 2014.
  21. 21.P.-L. Loh and M. J. Wainwright. Regularized m-estimators with nonconvexity : statistical and algorithmic theory for local optima. In arXiv:1305.2436v1, 2013.
  22. 22.R. Pascanu, T. Mikolov, and Y. Bengio. On the difficulty of training recurrent neural networks. In arXiv:1211.5063v2, 2014.
  23. 23.A. Rakhlin, O. Shamir, and K. Sridharan. Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes. In ICML, 2012.
  24. 24.J. Schmidhuber. Multi-column deep neural networks for image classification. In CVPR, 2012.
  25. 25.O. Shamir and T. Zhang. Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes. In ICML, 2013.
  26. 26.J. Snoek, R. P. Adams, and H. Larochelle. Nonparametric guidance of autoencoder representations using label information. J. of Machine Learning Research, 13:2567–2588, 2012.
  27. 27.N. Srivastava and R. Salakhutdinov. Discriminative transfer learning with tree-based priors. In NIPS, 2013.
  28. 28.Y. Tang. Deep learning using linear support vector machines. In Workshop on Representational Learning, ICML, 2013.
  29. 29.V. N. Vapnik. The nature of statistical learning theory. In Springer, New York, 1995.
  30. 30.J. Weston and F. Ratle. Deep learning via semi-supervised embedding. In ICML, 2008.
  31. 31.M. Zeiler and R. Fergus. Visualizing and understanding convolutional networks. In arXiv 1311.2901, 2013.
  32. 32.M. D. Zeiler and R. Fergus. Stochastic pooling for regularization of deep convolutional neural networks. In ICLR, 2013.

Citation

MLA
Lee, C.-Y., et al. “Deeply-Supervised Nets”. arXiv, 2014, https://doi.org/10.48550/arxiv.1409.5185.
APA
Lee, C.-Y., Xie, S., Gallagher, P., Zhang, Z., & Tu, Z. (2014). Deeply-Supervised Nets. arXiv. https://doi.org/10.48550/arxiv.1409.5185
Chicago
Lee, C.-Y., S. Xie, P. Gallagher, Z. Zhang, and Z. Tu. 2014. “Deeply-Supervised Nets”. Preprint, ArXiv. https://doi.org/10.48550/arxiv.1409.5185.
Harvard
Lee, C.-Y. et al. (2014) “Deeply-Supervised Nets”. arXiv. Available at: https://doi.org/10.48550/arxiv.1409.5185.
Vancouver
1. Lee C-Y, Xie S, Gallagher P, Zhang Z, Tu Z (2014) Deeply-Supervised Nets. https://doi.org/10.48550/arxiv.1409.5185

BibTeX

@misc{https://doi.org/10.48550/arxiv.1409.5185,
  doi = {10.48550/ARXIV.1409.5185},
  url = {https://arxiv.org/abs/1409.5185},
  author = {Lee, Chen-Yu and Xie, Saining and Gallagher, Patrick and Zhang, Zhengyou and Tu, Zhuowen},
  keywords = {Machine Learning (stat.ML), Computer Vision and Pattern Recognition (cs.CV), Machine Learning (cs.LG), Neural and Evolutionary Computing (cs.NE), FOS: Computer and information sciences, FOS: Computer and information sciences},
  title = {Deeply-Supervised Nets},
  publisher = {arXiv},
  year = {2014},
  copyright = {Creative Commons Attribution Non Commercial Share Alike 3.0 Unported}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: Published with permission