A Theoretical Analysis of Feature Pooling in Visual Recognition

Y-Lan BoureauJean PonceYann LeCun

article2010ICML1,409 citations

Establishes a rigorous theoretical framework explaining why max pooling outperforms average pooling for sparse visual features and demonstrates how pooling cardinality and dictionary size directly govern recognition accuracy.

Listen

Visual recognition systems often rely on spatial pooling—the process of aggregating nearby visual features into summary statistics—to achieve invariance to transformations, reduce noise, and compact data representations. While empirical studies show that the choice between average pooling and max pooling drastically affects recognition accuracy, the fundamental mechanisms governing their performance have remained poorly understood. Prior analyses have been clouded by confounding factors, particularly the relationship between feature sparsity, pooling neighborhood size, and feature extraction resolution.

The article establishes a theoretical framework to explain and quantify how pooling operations impact class separability in visual categorization. It evaluates the statistical behavior of average and max pooling across binary and continuous sparse representations, demonstrating how pooling choices can be systematically optimized.

To evaluate these dynamics, the authors developed statistical models analyzing class separability under independent and identically distributed Bernoulli, Gaussian, and exponential feature distributions. They then validated these theoretical predictions through empirical object and scene recognition experiments on two standard benchmark datasets (the 15 Scenes and Caltech-101 benchmarks), evaluating linear classification across varying codebook sizes, pooling cardinalities, and smoothing strategies.

The investigation produced several key findings regarding visual feature aggregation. First, max pooling proves mathematically superior to average pooling primarily when dealing with highly sparse features (features with very low probabilities of activation), whereas average pooling excels when features are dense or frequent. Second, for binary features, taking the maximum over all available samples in a region is suboptimal; instead, accuracy peaks at intermediate pooling cardinalities that scale with the size of the feature dictionary. Third, applying statistical smoothing—either by averaging max-pooled estimates over smaller subsets or by applying a theoretical expectation formula directly to the mean feature response—consistently outperforms standard max pooling and average pooling across all tested dictionary sizes. Finally, unified continuous formulations, such as the normalized vector norm, bridge average and max pooling and deliver robust performance across configurations.

These findings indicate that system designers do not need to rely on complex, non-linear classification kernels or computationally intensive pretraining to achieve high recognition accuracy. Properly calibrating the pooling step allows standard linear classifiers to match or exceed the performance of more complex architectures. This insight provides a direct pathway to reduce computational overhead, optimize training and inference speeds, and improve model robustness across diverse vision pipelines.

Based on these results, engineering teams should avoid default, full-cardinality max pooling when using vector-quantized binary features. Instead, systems should implement smoothed estimators—specifically applying the closed-form expectation transform to average-pooled responses—and adjust effective pooling cardinalities to match dictionary sparsity. For continuous sparse codes, systems should retain larger cardinalities or tune continuous norm-based pooling functions to find optimal operating points. Further research should extend this framework to adapt pooling parameters dynamically per individual feature type.

The primary limitation of this study lies in its foundational assumption of feature independence, which does not fully capture the strong spatial correlations present in real-world images. Additionally, the simplified analytical models only partially predict empirical behavior for continuous sparse codes, where complex mixture distributions arise. Nevertheless, the broad consistency between the theoretical models and experimental benchmarks provides high confidence in the recommended pooling optimizations for practical visual recognition systems.

Boureau et al (2010).pdf
  • Paper: Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition, Kaiming He et al. (2014). This work extends spatial pooling principles into deep convolutional networks via spatial pyramid pooling, enabling variable-sized input handling and robust multi-level feature aggregation.
  • Paper: Fine-Tuning CNN Image Retrieval with No Human Annotation, Filip Radenovic et al. (2017). This paper builds on the comparative mechanics of max and average pooling by introducing a generalized-mean (GeM) pooling layer that learns the optimal pooling parameter for image retrieval.
  • Paper: Striving for Simplicity: The All Convolutional Net, Jost Tobias Springenberg et al. (2014). This study critically re-evaluates the necessity of spatial max pooling in deep convolutional architectures, demonstrating that strided convolutions can effectively replace dedicated pooling operations.
  • Paper: Learning Deep Features for Discriminative Localization, Bolei Zhou et al. (2016). This work utilizes global average pooling at the network's terminal layer to retain spatial localization information and generate class activation maps for visual interpretability.
  • Paper: NetVLAD: CNN Architecture for Weakly Supervised Place Recognition, Relja Arandjelović et al. (2015). This paper advances feature pooling in deep architectures by introducing NetVLAD, a differentiable aggregation layer that extends traditional local feature pooling for place recognition.
  • Paper: MetaFormer is Actually What You Need for Vision, Weihao Yu et al. (2021). This paper evaluates the fundamental efficacy of basic spatial pooling by showing that replacing complex token-mixing attention mechanisms with simple average pooling yields highly competitive vision architectures.
Cover for A Theoretical Analysis of Feature Pooling in Visual Recognition

Abstract

Many modern visual recognition algorithms incorporate a step of spatial 'pooling', where the outputs of several nearby feature detectors are combined into a local or global 'bag of features', in a way that preserves task-related information while removing irrelevant details. Pooling is used to achieve invariance to image transformations, more compact representations, and better robustness to noise and clutter. Several papers have shown that the details of the pooling operation can greatly influence the performance, but studies have so far been purely empirical. In this paper, we show that the reasons underlying the performance of various pooling methods are obscured by several confounding factors, such as the link between the sample cardinality in a spatial pool and the resolution at which low-level features have been extracted. We provide a detailed theoretical analysis of max pooling and average pooling, and give extensive empirical comparisons for object recognition tasks.

Table of Contents

  • 1. Introduction
  • 2. Pooling Binary Features
  • 2.1. Model
  • 2.2. Distribution Separability
  • 2.2.1. Means' Separation of Max-Pooled Features
  • 2.2.2. Variance of Max-Pooled Features
  • 2.2.3. Conclusions and Predictions
  • 2.3. Experiments
  • 2.3.1. Optimal Pooling Cardinality
  • 2.3.2. Combining Multiple Pooling Cardinalities
  • 2.3.3. Practical Consequences
  • 3. Pooling Continuous Sparse Codes
  • 3.1. Influence of Pooling Cardinality
  • 3.2. Experiments
  • 3.3. Mixture Distribution
  • 4. Transition from Average to Max Pooling
  • 5. Discussion
  • References

Knowls

  1. Knowl 1 — Class Separability and Variance of Max Pooling for Binary Features

    theoretical result

    Let a feature response across PP spatial pooling locations be modeled as an independent and identically distributed (i.i.d.) Bernoulli random variable vector v=(v1,…,vP)∈{0,1}Pv = (v_1, \dots, v_P) \in \{0, 1\}^P with activation probability α=P(vi=1)\alpha = \mathbb{P}(v_i = 1). The max-pooled feature fm(v)=max⁡i=1,…,Pvif_m(v) = \max_{i=1,\dots,P} v_i is a Bernoulli random variable with expected value μm\mu_m and variance σm2\sigma_m^2 given by:

    μm=1−(1−α)P\mu_m = 1 - (1 - \alpha)^P

    σm2=(1−(1−α)P)(1−α)P\sigma_m^2 = (1 - (1 - \alpha)^P)(1 - \alpha)^P

    For two classes C1C_1 and C2C_2 with class-conditional activation probabilities α1=P(vi=1∣C1)\alpha_1 = \mathbb{P}(v_i = 1 \mid C_1) and α2=P(vi=1∣C2)\alpha_2 = \mathbb{P}(v_i = 1 \mid C_2), the separation distance between class-conditional expectations ϕ(P)=∣E(fm∣C1)−E(fm∣C2)∣=∣(1−α2)P−(1−α1)P∣\phi(P) = |\mathbb{E}(f_m \mid C_1) - \mathbb{E}(f_m \mid C_2)| = |(1 - \alpha_2)^P - (1 - \alpha_1)^P| is strictly increasing on the interval [0,PM][0, P_M] and strictly decreasing on [PM,∞)[P_M, \infty), with lim⁡P→0ϕ(P)=0\lim_{P \to 0} \phi(P) = 0 and lim⁡P→∞ϕ(P)=0\lim_{P \to \infty} \phi(P) = 0. The peak cardinality PMP_M is:

    PM=∣log⁡(log⁡(1−α2)log⁡(1−α1))log⁡(1−α11−α2)∣P_M = \left| \frac{\log\left(\frac{\log(1 - \alpha_2)}{\log(1 - \alpha_1)}\right)}{\log\left(\frac{1 - \alpha_1}{1 - \alpha_2}\right)} \right|

    The variance σm2\sigma_m^2 attains its maximum value of 0.50.5 at pool cardinality P=log⁡2∣log⁡(1−α)∣P = \frac{\log 2}{|\log(1 - \alpha)|} and approaches 00 as P→0P \to 0 or P→∞P \to \infty.

  2. Knowl 2 — Comparison of Max and Average Pooling for Sparse Binary Codes

    theoretical result

    For an i.i.d. Bernoulli feature vi∈{0,1}v_i \in \{0, 1\} with class-conditional probabilities α1=P(vi=1∣C1)\alpha_1 = \mathbb{P}(v_i = 1 \mid C_1) and α2=P(vi=1∣C2)\alpha_2 = \mathbb{P}(v_i = 1 \mid C_2) (assuming α1>α2\alpha_1 > \alpha_2), average pooling fa(v)=1P∑i=1Pvif_a(v) = \frac{1}{P}\sum_{i=1}^P v_i yields an expected separation distance ϕa=∣α1−α2∣\phi_a = |\alpha_1 - \alpha_2| that is constant with respect to pool cardinality PP, while its variance σa2=α(1−α)/P\sigma_a^2 = \alpha(1 - \alpha)/P decreases monotonically at rate 1/P1/P.

    In contrast, max pooling yields a larger difference between class expectations than average pooling (i.e., ϕ(P)>ϕa\phi(P) > \phi_a) over a range of cardinalities if and only if the peak cardinality satisfies PM>1P_M > 1. The condition PM≤1P_M \le 1 implies α1>1−1/e≈0.632\alpha_1 > 1 - 1/e \approx 0.632, which does not occur in standard visual codebooks (e.g., kk-word codebooks where the average activation is 1/k≪0.6321/k \ll 0.632).

    Max pooling is especially advantageous in the sparse regime α2≪α1≪1\alpha_2 \ll \alpha_1 \ll 1, where the two classes enter their asymptotic regimes (near-zero variance, high activation probability) at widely different cardinalities 1/α11/\alpha_1 and 1/α21/\alpha_2. The characteristic pool cardinality scale of the transition to the asymptotic regime grows proportionally to the dictionary size kk.

  3. Knowl 3 — Smooth Maximum Pooling Operator

    model/method

    Smooth maximum pooling computes the theoretical expectation of the maximum of PP Bernoulli random variables, using the empirical average activation α=1N∑i=1Nvi\alpha = \frac{1}{N} \sum_{i=1}^N v_i over all NN available patch samples in a spatial region as input:

    fSM(α;P)=1−(1−α)Pf_{\text{SM}}(\alpha; P) = 1 - (1 - \alpha)^P

    Here, P>0P > 0 is a continuous hyperparameter representing the effective pooling cardinality. This formulation replaces the non-smooth, sample-dependent maximum with an analytical nonlinear transformation applied to the average-pooled feature α\alpha.

    This method allows the effective pooling cardinality PP to be tuned independently of the physical number of patch samples NN available in a spatial pooling region, avoiding sample constraints when codebooks are large while smoothing out the high variance of a single sample maximum.

  4. Knowl 4 — Max Pooling Asymptotics for Exponentially Distributed Continuous Sparse Features

    theoretical result

    Let continuous sparse feature responses x1,…,xPx_1, \dots, x_P be i.i.d. random variables drawn from an exponential distribution with rate parameter λ>0\lambda > 0, having probability density function p(x)=λe−λxp(x) = \lambda e^{-\lambda x} for x≥0x \ge 0, mean 1/λ1/\lambda, and variance 1/λ21/\lambda^2. The max-pooled statistic fm=max⁡i=1,…,Pxif_m = \max_{i=1,\dots,P} x_i has cumulative distribution function Fm(x)=(1−e−λx)PF_m(x) = (1 - e^{-\lambda x})^P, with exact mean μm\mu_m and variance σm2\sigma_m^2 given by:

    μm=H(P)λ\mu_m = \frac{H(P)}{\lambda}

    σm2=1λ2∑l=1P1l(2H(l)−H(P))\sigma_m^2 = \frac{1}{\lambda^2} \sum_{l=1}^P \frac{1}{l} \left( 2H(l) - H(P) \right)

    where H(P)=∑i=1P1iH(P) = \sum_{i=1}^P \frac{1}{i} is the PP-th harmonic number.

    As P→∞P \to \infty, using the asymptotic expansion H(P)=log⁡P+γ+o(1)H(P) = \log P + \gamma + o(1) (where γ\gamma is Euler's constant), the class-conditional mean separation between rates λ1\lambda_1 and λ2\lambda_2 grows as O(log⁡P)O(\log P), whereas the standard deviation grows as O(log⁡P)O(\sqrt{\log P}). Consequently, the signal-to-noise ratio of raw max pooling over continuous exponential features grows monotonically with pool cardinality PP, but at a slower rate than the O(P)O(\sqrt{P}) noise reduction achieved by smoothing or averaging.

  5. Knowl 5 — Continuous Parametric Families Interpolating Average and Max Pooling

    model/method

    Four continuous parametric families bridge the continuum from average pooling to max pooling over a set of NN feature activations v=(v1,…,vN)v = (v_1, \dots, v_N):

    1. Normalized PP-norm Pooling: fP(v)=(1N∑i=1NviP)1/Pf_P(v) = \left( \frac{1}{N} \sum_{i=1}^N v_i^P \right)^{1/P} Yields average pooling at P=1P = 1, square-root pooling at P=2P = 2, and exact max pooling as P→∞P \to \infty.

    2. Softmax-Weighted Average: fsoftmax(v;β)=∑i=1N(exp⁡(βvi)∑j=1Nexp⁡(βvj))vif_{\text{softmax}}(v; \beta) = \sum_{i=1}^N \left( \frac{\exp(\beta v_i)}{\sum_{j=1}^N \exp(\beta v_j)} \right) v_i Yields average pooling at β=0\beta = 0 and max pooling as β→∞\beta \to \infty.

    3. Log-Sum-Exp Pooling: flogExp(v;β)=1βlog⁡(1N∑i=1Nexp⁡(βvi))f_{\text{logExp}}(v; \beta) = \frac{1}{\beta} \log \left( \frac{1}{N} \sum_{i=1}^N \exp(\beta v_i) \right) Yields average pooling as β→0\beta \to 0 and max pooling as β→∞\beta \to \infty.

    4. Expected Maximum Transformation: fmaxExp(α;P)=1−(1−α)Pf_{\text{maxExp}}(\alpha; P) = 1 - (1 - \alpha)^P where α=1N∑i=1Nvi\alpha = \frac{1}{N} \sum_{i=1}^N v_i is the empirical mean activation over the pool.

  6. Knowl 6 — Mixture Distribution Model for Continuous Sparse Feature Separability

    model/method

    To model continuous sparse image codes where most coordinates are exactly zero, feature activations are represented as a mixture distribution of a Dirac delta function δ(0)\delta(0) and an exponential distribution:

    p(x)=(1−w)δ(x)+wλe−λx,x≥0p(x) = (1 - w) \delta(x) + w \lambda e^{-\lambda x}, \quad x \ge 0

    where w∈[0,1]w \in [0, 1] is the activation probability and λ\lambda is the scale parameter of the active component.

    When the mixture weight ww fluctuates across images due to background clutter or unrelated content, the image-level empirical average activation shifts between 00 and w/λw/\lambda, introducing substantial inter-image variance and overlapping class-conditional distributions under average pooling. Max pooling is less sensitive to shifts in ww because it captures the presence of rare high-magnitude activations from the exponential tail, explaining why max pooling outperforms average pooling on sparse continuous representations.

  7. Knowl 7 — Spatial Pyramid Classification Accuracy with Smooth Maximum Pooling

    data/table

    The table below compares image classification accuracy (percentage correct ±\pm standard deviation over 10 splits) of standard max pooling (Max) versus smooth maximum pooling (SM) using binary features (1-of-kk vector quantized SIFT) across three codebook sizes (k∈{256,512,1024}k \in \{256, 512, 1024\}) in a 3-level spatial pyramid with a linear one-versus-all SVM on the Caltech-101 and 15 Scenes datasets. Smooth maximum pooling uses P=256P = 256 for k∈{256,512}k \in \{256, 512\} and P=512P = 512 for k=1024k = 1024.

    Dataset Method k=256k = 256 k=512k = 512 k=1024k = 1024
    Caltech 101 Max 67.5±1.067.5 \pm 1.0 69.2±1.169.2 \pm 1.1 71.0±0.871.0 \pm 0.8
    Caltech 101 SM 68.6±0.968.6 \pm 0.9 70.0±1.270.0 \pm 1.2 71.8±0.871.8 \pm 0.8
    15 Scenes Max 77.9±0.777.9 \pm 0.7 79.4±0.579.4 \pm 0.5 80.2±0.480.2 \pm 0.4
    15 Scenes SM 78.2±0.478.2 \pm 0.4 79.9±0.579.9 \pm 0.5 80.5±0.680.5 \pm 0.6

    Smooth maximum pooling consistently improves classification performance over standard max pooling across all codebook sizes on both object (Caltech-101) and scene (15 Scenes) benchmarks.

  8. Knowl 8 — Effect of Pool Cardinality and Smoothing Across Feature Types

    empirical result

    Empirical evaluation on Caltech-101 and 15 Scenes reveals distinct pooling behaviors depending on whether features are binary or continuous sparse codes:

    1. Binary Features:

      • For unsmoothed max pooling with a single estimate, recognition accuracy exhibits a peak at an intermediate pooling cardinality PP when codebooks are small (k=128,256k = 128, 256). This optimal cardinality shifts towards larger values as codebook size kk increases.
      • Smoothing the max-pooled estimate (either via an empirical average of maxima across subsamples or via the theoretical expectation 1−(1−α)P1 - (1 - \alpha)^P) consistently improves performance over raw max pooling and average pooling across all codebook sizes.
    2. Continuous Sparse Features:

      • For unsmoothed max pooling, recognition accuracy increases monotonically with pool cardinality PP, exhibiting no intermediate peak regardless of codebook size.
      • Smoothing the maximum over smaller sub-pools provides modest performance improvements, explaining part of the empirical advantage of multi-level spatial pyramids over single-pool architectures.
  9. Knowl 9 — Effect of Combining Multiple Pooling Cardinalities on Classification

    data/table

    The table below shows classification accuracy (percentage correct ±\pm standard deviation over 10 splits) on Caltech-101 and 15 Scenes for a codebook of size k=256k = 256 under whole-image pooling on binary codes. Features are evaluated using a single cardinality ('One') versus jointly concatenating features from multiple pooling cardinalities ('Joint', combining the smallest cardinality with 512 and 1024) across average pooling (Avg), standard max pooling (Max), and smooth maximum pooling (SM, 1−(1−α)P1 - (1 - \alpha)^P).

    Dataset Pooling Strategy Cardinality 1024 Cardinality 512 Cardinality 256
    Caltech 101 Avg, One 32.4±1.132.4 \pm 1.1 31.3±1.031.3 \pm 1.0 28.6±1.128.6 \pm 1.1
    Caltech 101 Avg, Joint - 31.9±1.231.9 \pm 1.2 32.1±1.232.1 \pm 1.2
    Caltech 101 Max, One 31.7±1.431.7 \pm 1.4 32.7±1.332.7 \pm 1.3 30.4±2.330.4 \pm 2.3
    Caltech 101 Max, Joint - 34.4±0.734.4 \pm 0.7 35.8±0.935.8 \pm 0.9
    Caltech 101 SM, One 37.9±0.637.9 \pm 0.6 40.5±0.740.5 \pm 0.7 42.0±1.442.0 \pm 1.4
    Caltech 101 SM, Joint - 39.4±1.339.4 \pm 1.3 40.6±0.840.6 \pm 0.8
    15 Scenes Avg, One 69.8±0.769.8 \pm 0.7 68.7±0.868.7 \pm 0.8 66.3±0.766.3 \pm 0.7
    15 Scenes Avg, Joint - 69.6±0.769.6 \pm 0.7 69.2±1.069.2 \pm 1.0
    15 Scenes Max, One 63.5±0.663.5 \pm 0.6 64.8±0.764.8 \pm 0.7 64.3±0.464.3 \pm 0.4
    15 Scenes Max, Joint - 65.4±0.665.4 \pm 0.6 67.1±0.667.1 \pm 0.6
    15 Scenes SM, One 67.2±0.867.2 \pm 0.8 70.4±0.770.4 \pm 0.7 72.6±0.772.6 \pm 0.7
    15 Scenes SM, Joint - 69.2±0.769.2 \pm 0.7 70.7±0.770.7 \pm 0.7

    Combining multiple pooling cardinalities improves accuracy for standard max pooling (e.g., from 30.4%30.4\% to 35.8%35.8\% on Caltech-101 at cardinality 256) because multi-scale pooling acts as an implicit smoothing mechanism. When explicit smoothing is applied via smooth max (SM), joint cardinalities provide no additional benefit over the single optimal cardinality.

Coverage note — None was omitted; all major theoretical models, closed-form derivations, parametric interpolation families, and experimental tables/findings were extracted.

References

  1. 1.Boureau, Y, Bach, F, LeCun, Y, and Ponce, J. Learning mid-level features for recognition. In CVPR, 2010.
  2. 2.Bruckstein, A and Cover, T. Monotonicity of linear separability under translation. IEEE TRANS. PATTERN ANAL. MACH. INTELLIG., 7(3):355–357, 1985.
  3. 3.Dalal, N and Triggs, B. Histograms of oriented gradients for human detection. In CVPR, 2005.
  4. 4.Fei-Fei, L, Fergus, R, and Perona, P. Learning generative visual models from few training examples. In CVPR Workshop GMBV, 2004.
  5. 5.Fukushima, K and Miyake, S. Neocognitron: A new algorithm for pattern recognition tolerant of deformations and shifts in position. Pattern Recognition, 1982.
  6. 6.Hubel, D. H and Wiesel, T. N. Receptive fields, binocular interaction and functional architecture in the cat’s visual cortex. J Physiol, 160:106–154, Jan 1962.
  7. 7.Jarrett, K, Kavukcuoglu, K, Ranzato, M, and LeCun, Y. What is the best multi-stage architecture for object recognition? In (ICCV’09). IEEE, 2009.
  8. 8.Koenderink, J and Van Doorn, A. The structure of locally orderless images. IJCV, 31(2/3):159–168, 1999.
  9. 9.Lazebnik, S, Schmid, C, and Ponce, J. Beyond bags of features: Spatial pyramid matching for recognizing natural scene categories. In CVPR, 2006.
  10. 10.LeCun, Y, Boser, B, Denker, J. S, Henderson, D, Howard, R. E, Hubbard, W, and Jackel, L. D. Handwritten digit recognition with a back-propagation network. In NIPS, 1989.
  11. 11.LeCun, Y, Bottou, L, Bengio, Y, and Haffner, P. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998.
  12. 12.Lowe, D. Distinctive image features from scale-invariant keypoints. Int. J. of Comp. Vision, 60(4):91–110, 2004.
  13. 13.Pinto, N, Cox, D, and DiCarlo, J. Why is real-world visual object recognition hard. PLoS Computational Biology, 4(1):151–156, 2008.
  14. 14.Ranzato, M, Boureau, Y, and LeCun, Y. Sparse feature learning for deep belief networks. In NIPS, 2007.
  15. 15.Serre, T, Wolf, L, and Poggio, T. Object recognition with features inspired by visual cortex. In CVPR, 2005.
  16. 16.Sivic, J and Zisserman, A. Video Google: A text retrieval approach to object matching in videos. In ICCV, 2003. URL http://www.robots.ox.ac.uk/˜vgg.
  17. 17.Yang, J, Yu, K, Gong, Y, and Huang, T. Linear Spatial Pyramid Matching Using Sparse Coding for Image Classification. In CVPR, 2009.
  18. 18.Zhang, J, Marszalek, M, Lazebnik, S, and Schmid, C. Local features and kernels for classifcation of texture and object categories: An in-depth study. Int. J. of Comp. Vision, 73(2):213–238, 2007.

Citation

MLA
Boureau, Y.-L., et al. “A Theoretical Analysis of Feature Pooling in Visual Recognition”. 2010, http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.170.864.
APA
Boureau, Y.-L., Ponce, J., & LeCun, Y. (2010). A Theoretical Analysis of Feature Pooling in Visual Recognition. http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.170.864
Chicago
Boureau, Y.-L., J. Ponce, and Y. LeCun. 2010. “A Theoretical Analysis of Feature Pooling in Visual Recognition”. Preprint. http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.170.864.
Harvard
Boureau, Y.-L., Ponce, J. and LeCun, Y. (2010) “A Theoretical Analysis of Feature Pooling in Visual Recognition”. Available at: http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.170.864.
Vancouver
1. Boureau Y-L, Ponce J, LeCun Y (2010) A Theoretical Analysis of Feature Pooling in Visual Recognition.

BibTeX

@article{boureau2010theoretical,
  title = {A Theoretical Analysis of Feature Pooling in Visual Recognition},
  author = {Boureau, Y-Lan and Ponce, Jean and LeCun, Yann},
  year = {2010},
  url = {http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.170.864}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors