TextonBoost: Joint Appearance, Shape and Context Modeling for Multi-class Object Recognition and Segmentation

J. ShottonJ. WinnC. RotherA. Criminisi

article2006ECCV1,397 citations

Introduces the TextonBoost framework, which unifies appearance, shape, and spatial context using boosted texton features within a conditional random field to achieve accurate, multi-class semantic segmentation and object recognition.

Listen

Automatic visual recognition and pixel-level segmentation in digital images remain fundamental challenges in computer vision, particularly when handling numerous object categories under diverse real-world conditions such as varying lighting, viewpoints, and occlusions. Existing systems frequently separate detection from segmentation or rely on computationally prohibitive statistical modeling, preventing effective scaling to larger datasets and complex environments. The article demonstrates a unified framework that simultaneously performs object recognition and semantic segmentation across multiple classes efficiently.

The evaluated approach combines machine learning classification and probabilistic graphical models by integrating texture-based shape features into a conditional random field. The system generates texton maps to capture texture and pairs them with spatial filters to jointly represent object shape, appearance, and visual context. To enable scalable execution across large databases, the training uses a shared boosting method with randomized feature evaluation, sub-sampling, and piecewise optimization. The full framework integrates shape-texture predictions with color, pixel location, and contrast-sensitive edge cues, resolving final pixel assignments via efficient graph-cut optimization. Performance was validated across three datasets, including a complex 21-class benchmark of natural photographs.

The findings establish that the unified approach achieves 72.2% overall pixel-level accuracy across 21 diverse object categories, outperforming random chance by approximately fifteen times. While the standalone boosted classifier achieved 69.6% accuracy, integrating edge, color, and location potentials raised accuracy by 2.6% and dramatically improved visual contour delineation and boundary sharpness. Furthermore, randomized feature selection accelerated the boosting training phase by roughly two orders of magnitude with minimal loss in accuracy. On benchmark 7-class datasets, the system delivered accuracy comparable to existing models—achieving 88.6% on the Sowerby dataset and 74.6% on Corel—while drastically reducing training and inference times by avoiding computationally expensive sampling methods.

These results demonstrate that joint modeling of shape, texture, and contextual layout can deliver highly competitive semantic segmentation while maintaining practical computational efficiency. By substantially lowering training runtimes from thousands of hours to manageable operational windows, this methodology enables practical deployment in systems processing massive image libraries. However, performance variations highlight that classes with high visual diversity or limited sample sizes, such as boats and chairs, exhibit lower accuracy, and structured objects are occasionally confused when sharing contextual backgrounds. The article recommends expanding training sample sizes, integrating explicit part-based detection for structured objects, and incorporating higher-level semantic scene context to further enhance overall classification reliability.

Cover for TextonBoost: Joint Appearance, Shape and Context Modeling for Multi-class Object Recognition and Segmentation

Abstract

This paper proposes a new approach to learning a discriminative model of object classes, incorporating appearance, shape and context information efficiently. The learned model is used for automatic visual recognition and semantic segmentation of photographs. Our discriminative model exploits novel features, based on textons, which jointly model shape and texture. Unary classification and feature selection is achieved using shared boosting to give an efficient classifier which can be applied to a large number of classes. Accurate image segmentation is achieved by incorporating these classifiers in a conditional random field. Efficient training of the model on very large datasets is achieved by exploiting both random feature selection and piecewise training methods.

High classification and segmentation accuracy are demonstrated on three different databases: i) our own 21-object class database of photographs of real objects viewed under general lighting conditions, poses and viewpoints, ii) the 7-class Corel subset and iii) the 7-class Sowerby database used in [1]. The proposed algorithm gives competitive results both for highly textured (e.g. grass, trees), highly structured (e.g. cars, faces, bikes, aeroplanes) and articulated objects (e.g. body, cow).

Table of Contents

  • 1 Introduction
  • 2 Image Databases
  • 3 A Conditional Random Field Model of Object Classes
  • 3.1 Learning the CRF Parameters
  • 3.2 Inference in the CRF Model
  • 4 Boosted Learning of Shape, Texture and Context
  • 5 Results and Comparisons
  • 6 Conclusions
  • References

Knowls

  1. Knowl 1 — Conditional Random Field Formulation for Multi-Class Object Segmentation

    model/method

    The conditional random field (CRF) models the conditional distribution over a 2D pixel labeling c={ci}i∈V\mathbf{c} = \{c_i\}_{i \in \mathcal{V}} (where ci∈Cc_i \in \mathcal{C} is a discrete class label from a set of object classes C\mathcal{C}) given an observed input image x\mathbf{x} as:

    log⁡P(c∣x,θ)=∑i∈Vψi(ci,x;θψ)+∑i∈Vπ(ci,xi;θπ)+∑i∈Vλ(ci,i;θλ)+∑(i,j)∈Eϕ(ci,cj,gij(x);θϕ)−log⁡Z(θ,x)\log P(\mathbf{c} \mid \mathbf{x}, \boldsymbol{\theta}) = \sum_{i \in \mathcal{V}} \psi_i(c_i, \mathbf{x}; \boldsymbol{\theta}_\psi) + \sum_{i \in \mathcal{V}} \pi(c_i, \mathbf{x}_i; \boldsymbol{\theta}_\pi) + \sum_{i \in \mathcal{V}} \lambda(c_i, i; \boldsymbol{\theta}_\lambda) + \sum_{(i,j) \in \mathcal{E}} \phi(c_i, c_j, \mathbf{g}_{ij}(\mathbf{x}); \boldsymbol{\theta}_\phi) - \log Z(\boldsymbol{\theta}, \mathbf{x})

    where:

    • V\mathcal{V} is the set of pixel nodes in the image.
    • E\mathcal{E} is the set of edges connecting adjacent pixels in a 4-connected grid.
    • Z(θ,x)Z(\boldsymbol{\theta}, \mathbf{x}) is the partition function summing over all possible label configurations.
    • θ={θψ,θπ,θλ,θϕ}\boldsymbol{\theta} = \{\boldsymbol{\theta}_\psi, \boldsymbol{\theta}_\pi, \boldsymbol{\theta}_\lambda, \boldsymbol{\theta}_\phi\} denotes the complete set of model parameters.
    • ψi(ci,x;θψ)\psi_i(c_i, \mathbf{x}; \boldsymbol{\theta}_\psi) is the shape-texture unary potential capturing appearance, shape, and spatial context.
    • π(ci,xi;θπ)\pi(c_i, \mathbf{x}_i; \boldsymbol{\theta}_\pi) is the unary color potential capturing image-specific color distributions.
    • λ(ci,i;θλ)\lambda(c_i, i; \boldsymbol{\theta}_\lambda) is the unary location potential capturing absolute spatial position priors.
    • ϕ(ci,cj,gij(x);θϕ)\phi(c_i, c_j, \mathbf{g}_{ij}(\mathbf{x}); \boldsymbol{\theta}_\phi) is the pairwise edge potential penalizing label discontinuities modulated by image contrast.
  2. Knowl 2 — Shape Filter Features for Joint Texture, Shape, and Context Representation

    definition

    A shape filter is a spatial feature defined over an image texton map. The texton map is constructed by convolving an image with a 17-dimensional filter bank (scaled Gaussians on CIELab color channels, plus first xx- and yy-derivatives and Laplacians of Gaussians on luminance) and assigning each pixel to the nearest cluster centroid determined by KK-means clustering under Mahalanobis distance.

    A shape filter is parameterized by a texton index t∈{1,…,K}t \in \{1, \dots, K\} and a rectangular bounding region rr defined relative to an offset origin. The feature response v(i,r,t)v(i, r, t) at pixel location ii is defined as the total count of pixels with texton tt lying within the rectangle rr offset from ii:

    v(i,r,t)=∑j∈r(i)[T(j)=t]v(i, r, t) = \sum_{j \in r(i)} [T(j) = t]

    where r(i)r(i) denotes the rectangle rr shifted by pixel position ii, T(j)T(j) is the texton index at pixel jj, and [⋅][\cdot] is the Iverson bracket. By leveraging integral images over texton indicator maps (KK integral images per image), v(i,r,t)v(i, r, t) is evaluated in O(1)O(1) constant time per pixel. Depending on the size and spatial offset of rr, shape filters capture local texture (small offset), object shape (intermediate offset across object parts), or background context (large offset to surrounding scene elements).

  3. Knowl 3 — Joint Boosted Classifier for Shape-Texture Potentials

    model/method

    The unary shape-texture potential ψi(ci,x;θψ)\psi_i(c_i, \mathbf{x}; \boldsymbol{\theta}_\psi) is defined as the log-probability under a multi-class boosted classifier:

    ψi(ci,x;θψ)=log⁡P~i(ci∣x)\psi_i(c_i, \mathbf{x}; \boldsymbol{\theta}_\psi) = \log \tilde{P}_i(c_i \mid \mathbf{x})

    where P~i(ci∣x)\tilde{P}_i(c_i \mid \mathbf{x}) is obtained by a softmax transformation over an additive ensemble of MM weak classifiers hmh_m:

    P~i(ci∣x)=exp⁡(H(ci))∑c′∈Cexp⁡(H(c′)),H(c)=∑m=1Mhm(c)\tilde{P}_i(c_i \mid \mathbf{x}) = \frac{\exp(H(c_i))}{\sum_{c' \in \mathcal{C}} \exp(H(c'))}, \quad H(c) = \sum_{m=1}^M h_m(c)

    Each weak classifier h(c)h(c) is a decision stump shared across a subset of classes N⊆C\mathcal{N} \subseteq \mathcal{C}:

    h(ci)={a[v(i,r,t)>θ]+bif ci∈Nkciotherwiseh(c_i) = \begin{cases} a [v(i, r, t) > \theta] + b & \text{if } c_i \in \mathcal{N} \\ k_{c_i} & \text{otherwise} \end{cases}

    where v(i,r,t)v(i, r, t) is a shape filter response, θ∈Θ\theta \in \Theta is a scalar threshold, aa and bb are regression parameters for sharing classes, and kcik_{c_i} is a class-specific constant handling class asymmetry. Feature sharing across class subsets N\mathcal{N} enables multi-class discrimination with evaluation cost sub-linear in the total number of classes ∣C∣|\mathcal{C}|, with optimal sharing subsets approximated greedily.

  4. Knowl 4 — Contrast-Sensitive Pairwise Edge Potential

    model/method

    The pairwise edge potential ϕ(ci,cj,gij(x);θϕ)\phi(c_i, c_j, \mathbf{g}_{ij}(\mathbf{x}); \boldsymbol{\theta}_\phi) is formulated as a contrast-sensitive Potts model defined on neighboring pixels (i,j)∈E(i, j) \in \mathcal{E}:

    ϕ(ci,cj,gij(x);θϕ)=−θϕTgij(x)[ci≠cj]\phi(c_i, c_j, \mathbf{g}_{ij}(\mathbf{x}); \boldsymbol{\theta}_\phi) = -\boldsymbol{\theta}_\phi^T \mathbf{g}_{ij}(\mathbf{x}) [c_i \neq c_j]

    where [ci≠cj]=1[c_i \neq c_j] = 1 if ci≠cjc_i \neq c_j and 00 otherwise. The edge feature vector gij(x)∈R2\mathbf{g}_{ij}(\mathbf{x}) \in \mathbb{R}^2 is:

    gij(x)=[exp⁡(−β∥xi−xj∥2)1]\mathbf{g}_{ij}(\mathbf{x}) = \begin{bmatrix} \exp(-\beta \|\mathbf{x}_i - \mathbf{x}_j\|^2) \\ 1 \end{bmatrix}

    where xi\mathbf{x}_i and xj\mathbf{x}_j are 3-dimensional color vectors (in CIELab color space) of adjacent pixels ii and jj. The parameter β\beta is computed per image as:

    β=12⟨∥xi−xj∥2⟩\beta = \frac{1}{2 \langle \|\mathbf{x}_i - \mathbf{x}_j\|^2 \rangle}

    where ⟨⋅⟩\langle \cdot \rangle denotes the average squared Euclidean color distance across all adjacent pixel pairs in the image. The unit element in gij\mathbf{g}_{ij} provides a learned constant bias that penalizes boundary length and suppresses small, isolated regions, weighted by the non-negative parameter vector θϕ=[θϕ,1,θϕ,2]T\boldsymbol{\theta}_\phi = [\theta_{\phi, 1}, \theta_{\phi, 2}]^T.

  5. Knowl 5 — Image-Specific Shared Gaussian Mixture Color Potential

    model/method

    The color potential π(ci,xi;θπ)\pi(c_i, \mathbf{x}_i; \boldsymbol{\theta}_\pi) captures image-specific color distributions for class instances present in an individual photograph:

    π(ci,xi;θπ)=log⁡∑k=1Kθπ(ci,k)P(k∣xi)\pi(c_i, \mathbf{x}_i; \boldsymbol{\theta}_\pi) = \log \sum_{k=1}^K \boldsymbol{\theta}_\pi(c_i, k) P(k \mid \mathbf{x}_i)

    where k∈{1,…,K}k \in \{1, \dots, K\} indexes Gaussian mixture model (GMM) components with means xˉk\bar{\mathbf{x}}_k and covariances Σk\boldsymbol{\Sigma}_k shared across all classes, such that P(x∣c)=∑kP(k∣c)N(x∣xˉk,Σk)P(\mathbf{x} \mid c) = \sum_k P(k \mid c) \mathcal{N}(\mathbf{x} \mid \bar{\mathbf{x}}_k, \boldsymbol{\Sigma}_k). Each pixel xi\mathbf{x}_i is assigned a fixed soft responsibility P(k∣xi)P(k \mid \mathbf{x}_i).

    At test time, the image-specific mixing parameters θπ\boldsymbol{\theta}_\pi are learned iteratively via Iterative Conditional Modes (ICM) based on the inferred segmentation c∗\mathbf{c}^*:

    θπ(ci,k)=(∑j[cj∗=ci]P(k∣xj)+απ∑jP(k∣xj)+απ)wπ\boldsymbol{\theta}_\pi(c_i, k) = \left( \frac{\sum_{j} [c_j^* = c_i] P(k \mid \mathbf{x}_j) + \alpha_\pi}{\sum_{j} P(k \mid \mathbf{x}_j) + \alpha_\pi} \right)^{w_\pi}

    where απ=0.1\alpha_\pi = 0.1 is a Dirichlet prior hyperparameter, and wπ=3w_\pi = 3 is a power parameter compensating for overcounting under piecewise inference. In practice, K=15K = 15 color components and two ICM iterations are used.

  6. Knowl 6 — Normalized Location Prior Potential

    model/method

    The location potential λ(ci,i;θλ)\lambda(c_i, i; \boldsymbol{\theta}_\lambda) models absolute spatial priors of object classes across the image canvas:

    λ(ci,i;θλ)=log⁡θλ(ci,i^)\lambda(c_i, i; \boldsymbol{\theta}_\lambda) = \log \boldsymbol{\theta}_\lambda(c_i, \hat{i})

    where i^\hat{i} is the 2D coordinate of pixel ii normalized onto a canonical reference square, ensuring scale invariance across varying image dimensions.

    The location parameter table θλ(ci,i^)\boldsymbol{\theta}_\lambda(c_i, \hat{i}) is learned directly from training pixel counts:

    θλ(ci,i^)=(Nci,i^+αλNi^+αλ)wλ\boldsymbol{\theta}_\lambda(c_i, \hat{i}) = \left( \frac{N_{c_i, \hat{i}} + \alpha_\lambda}{N_{\hat{i}} + \alpha_\lambda} \right)^{w_\lambda}

    where Nci,i^N_{c_i, \hat{i}} is the count of training pixels labeled with class cic_i at normalized position i^\hat{i}, Ni^=∑cNc,i^N_{\hat{i}} = \sum_c N_{c, \hat{i}} is the total number of training pixels at that location, αλ=1\alpha_\lambda = 1 is a weak Dirichlet prior hyperparameter, and wλw_\lambda is a power parameter tuned discriminatively (e.g., wλ∈[0.1,4]w_\lambda \in [0.1, 4]) to correct for overcounting during piecewise model combination.

  7. Knowl 7 — Piecewise CRF Training and Graph-Cut Inference Algorithm

    algorithm

    To circumvent intractable partition function computations and avoid slow Gibbs sampling, TextonBoost trains model potentials independently via piecewise training and computes the MAP labeling at test time using α\alpha-expansion graph cuts:

    Input: Image x\mathbf{x}, shape-texture boosted classifier H(c)H(c), location table θλ\boldsymbol{\theta}_\lambda, edge weights θϕ\boldsymbol{\theta}_\phi, shared color components P(k∣x)P(k \mid \mathbf{x})
    Output: MAP segmentation labeling c∗=arg⁡max⁡cP(c∣x,θ)\mathbf{c}^* = \arg\max_\mathbf{c} P(\mathbf{c} \mid \mathbf{x}, \boldsymbol{\theta})
    Initialize class labeling pixel-wise to the unary mode:
      for each pixel i∈Vi \in \mathcal{V} do
        ci∗=arg⁡max⁡c∈C(ψi(c,x)+λ(c,i))c_i^* = \arg\max_{c \in \mathcal{C}} (\psi_i(c, \mathbf{x}) + \lambda(c, i))
      end for
    for iteration = 1 to 2 do
      Update image-specific color parameters θπ(c,k)\boldsymbol{\theta}_\pi(c, k) for all c∈C,k∈{1,…,K}c \in \mathcal{C}, k \in \{1,\dots,K\} using current labeling c∗\mathbf{c}^*
      Compute unary energies for all i∈V,c∈Ci \in \mathcal{V}, c \in \mathcal{C}:
        Ei(c)=−(ψi(c,x)+π(c,xi;θπ)+λ(c,i))E_i(c) = -(\psi_i(c, \mathbf{x}) + \pi(c, \mathbf{x}_i; \boldsymbol{\theta}_\pi) + \lambda(c, i))
      Compute pairwise energies for all (i,j)∈E,(ci,cj)∈C×C(i, j) \in \mathcal{E}, (c_i, c_j) \in \mathcal{C} \times \mathcal{C}:
        Eij(ci,cj)=θϕTgij(x)[ci≠cj]E_{ij}(c_i, c_j) = \boldsymbol{\theta}_\phi^T \mathbf{g}_{ij}(\mathbf{x}) [c_i \neq c_j]
      Update c∗=arg⁡min⁡c(∑iEi(ci)+∑(i,j)Eij(ci,cj))\mathbf{c}^* = \arg\min_\mathbf{c} \left( \sum_{i} E_i(c_i) + \sum_{(i,j)} E_{ij}(c_i, c_j) \right) via α\alpha-expansion graph cuts
    end for
    return c∗\mathbf{c}^*
  8. Knowl 8 — Boosting Acceleration via Spatial Sub-Sampling and Randomized Feature Selection

    model/method

    Training boosted shape-filter classifiers on dense multi-class pixel sets is accelerated using two approximations:

    1. Spatial Sub-sampling: Feature responses and training labels are extracted only on a subsampled Δ×Δ\Delta \times \Delta pixel grid (Δ=3\Delta = 3 for small images, Δ=5\Delta = 5 for 320×240320 \times 240 images). Full-resolution integral images are retained at test time to support dense per-pixel inference. The small degree of shift-invariance introduced by sub-sampling is subsequently resolved at test time by the CRF edge and color potentials.
    2. Random Feature Selection: At each round of JointBoost, only a random fraction τ=0.003\tau = 0.003 of all possible rectangle-texton feature pairs (r,t)(r, t) is evaluated rather than performing exhaustive search over all candidate pairs.

    On a 21-class dataset with 276 training images, randomized feature selection reduces training time from approximately 14,000 hours to 42 hours (for 5,000 boosting rounds on a 2.1 GHz CPU) with negligible increase in training error.

  9. Knowl 9 — Segmentation and Recognition Accuracy Benchmarks

    data/table

    Semantic segmentation and recognition performance of TextonBoost evaluated across the 7-class Sowerby dataset (96×6496 \times 64 pixels), the 7-class Corel subset (180×120180 \times 120 pixels), and the 21-class MSRC database (320×240320 \times 240 pixels):

    Model Pixel Accuracy (%) Speed (Train / Test)
    Sowerby Corel Sowerby Corel
    TextonBoost – Full CRF 88.6% 74.6% 5h / 10s 12h / 30s
    TextonBoost – Unary Classifier Only 85.6% 68.4% – –
    He et al. – Multiscale CRF [1] 89.5% 80.0% Gibbs Gibbs
    He et al. – Unary Classifier Only [1] 82.4% 66.9% – –

    On the 21-class dataset, TextonBoost achieves an overall pixel-wise classification accuracy of 72.2% (compared to 4.76% for random chance, 69.6% for boosted unary alone, and 70.3% for the CRF without color potentials). On pre-segmented regions from the 21-class dataset, TextonBoost achieves 70.5% recognition accuracy, outperforming visual dictionary Gaussian models (67.6%).

  10. Knowl 10 — Class-Wise Recognition Patterns and Error Modes in TextonBoost

    empirical result

    Row-normalized confusion analysis across the 21-class MSRC dataset demonstrates that TextonBoost achieves highest pixel accuracy on classes with low visual variability and abundant training pixels: grass (97.6%), book (91.9%), tree (86.3%), road (86.0%), and sky (82.6%). Accuracy is lowest on classes with high visual variability, articulation, or sparse training data: boat (6.6%), dog (19.2%), bird (19.4%), chair (15.4%), and sign (35.1%).

    Two dominant systematic error patterns occur:

    1. Contextual Co-occurrence Confusion: Objects consistently appearing on grass in the training set are frequently misclassified as grass (e.g., cow at 30.9%, sheep at 25.5%, chair at 24.8%), an effect partly driven by loose ground-truth annotation boundaries.
    2. Structural Ambiguity: Rigid man-made classes (aeroplanes at 21.5%, chairs at 20.6%, signs at 31.5%, boats at 25.1%) are frequently mislabeled as buildings due to shared rectangular textures, straight lines, and planar features.

Coverage note — None was omitted; all key contributions—the CRF formulation, shape filter feature definitions, joint boostingUnary classifier, pairwise/color/location potentials, piecewise training, inference via graph cuts, randomized feature selection, benchmark results, and confusion error analysis—have been extracted.

References

  1. 1.He, X., Zemel, R.S., Carreira-Perpiñán, M.A.: Multiscale conditional random fields for image labeling. Proc. of IEEE CVPR (2004)
  2. 2.Fergus, R., Perona, P., Zisserman, A.: Object class recognition by unsupervised scale-invariant learning. In: CVPR'03. Volume II. (2003) 264–271
  3. 3.Berg, A.C., Berg, T.L., Malik, J.: Shape matching and object recognition using low distortion correspondences. In: CVPR. (2005)
  4. 4.Winn, J., Criminisi, A., Minka, T.: Categorization by learned universal visual dictionary. Int. Conf. of Computer Vision (2005)
  5. 5.Kumar, S., Herbert, M.: Discriminative fields for modeling spatial dependencies in natural images. In: NIPS. (2004)
  6. 6.Borenstein, E., Sharon, E., Ullman, S.: Combining top-down and bottom-up segmentation. In: Proceedings IEEE workshop on Perceptual Organization in Computer Vision, CVPR 2004. (2004)
  7. 7.Winn, J., Jojic, N.: LOCUS: Learning Object Classes with Unsupervised Segmentation. Proc. of IEEE ICCV. (2005)
  8. 8.Kumar, P., Torr, P., Zisserman, A.: Obj cut. Proc. of IEEE CVPR. (2005)
  9. 9.Leibe, B., Schiele, B.: Interleaved object categorization and segmentation. In: BMVC'03. Volume II. (2003) 264–271
  10. 10.Duygulu, P., Barnard, K., de Freitas, N., Forsyth, D.: Object recognition as machine translation: Learning a lexicon for a fixed image vocabulary. ECCV (2002)
  11. 11.Tu, Z., Chen, X., Yuille, A.L., Zhu, S.: Image parsing: Unifying segmentation, detection, and recognition. In: CVPR. (2003)
  12. 12.Konishi, S., Yuille, A.L.: Statistical cues for domain specific image segmentation with performance analysis. In: CVPR. (2000)
  13. 13.Lafferty, J., McCallum, A., Pereira, F.: Conditional random fields: Probabilistic models for segmenting and labeling sequence data. In: ICML. (2001)
  14. 14.Boykov, Y., Jolly, M.P.: Interactive graph cuts for optimal boundary and region segmentation of objects in n-d images. Proc. of IEEE ICCV. (2001)
  15. 15.Rother, C., Kolmogorov, V., Blake, A.: Interactive foreground extraction using iterated graph cuts. ACM Transactions on Graphics (SIGGRAPH'04). (2004)
  16. 16.Sutton, C., McCallum, A.: Piecewise training of undirected models. In: 21st Conference on Uncertainty in Artificial Intelligence. (2005)
  17. 17.Leung, T., Malik, J.: Representing and recognizing the visual appearance of materials using three-dimensional textons. IJCV 43 (2001) 29–44
  18. 18.Varma, M., Zisserman, A.: A statistical approach to texture classification from single images. International Journal of Computer Vision: Special Issue on Texture Analysis and Synthesis 62 (2005) 61–81
  19. 19.Viola, P., Jones, M.: Rapid object detection using a boosted cascade of simple features. In: CVPR01. (2001) I:511–518
  20. 20.Belongie, S., Malik, J., Puzicha, J.: Shape matching and object recognition using shape contexts. PAMI 24 (2002) 509–522
  21. 21.Torralba, A., Murphy, K., Freeman, W.: Sharing features: efficient boosting procedures for multiclass object detection. Proc. of IEEE CVPR (2004) 762–769
  22. 22.Friedman, J., Hastie, T., Tibshirani, R.: Additive logistic regression: a statistical view of boosting. Technical report, Dept. of Statistics, Stanford University. (1998)
  23. 23.Baluja, S., Rowley, H.A.: Boosting sex identification performance. In: AAAI. (2005) 1508–1513
  24. 24.Kumar, S., Hebert, M.: A hierarchical field framework for unified context-based classification. In: ICCV05. (2005) II: 1284–1291

Citation

MLA
Shotton, J., et al. “TextonBoost: Joint Appearance, Shape and Context Modeling for Multi-class Object Recognition and Segmentation”. Lecture Notes in Computer Science, Springer Berlin Heidelberg, 2006, pp. 1–5, https://doi.org/10.1007/11744023_1.
APA
Shotton, J., Winn, J., Rother, C., & Criminisi, A. (2006). TextonBoost: Joint Appearance, Shape and Context Modeling for Multi-class Object Recognition and Segmentation. In Lecture Notes in Computer Science (pp. 1–15). Springer Berlin Heidelberg. https://doi.org/10.1007/11744023_1
Chicago
Shotton, J., J. Winn, C. Rother, and A. Criminisi. 2006. “TextonBoost: Joint Appearance, Shape and Context Modeling for Multi-class Object Recognition and Segmentation”. In Lecture Notes in Computer Science. Springer Berlin Heidelberg. https://doi.org/10.1007/11744023_1.
Harvard
Shotton, J. et al. (2006) “TextonBoost: Joint Appearance, Shape and Context Modeling for Multi-class Object Recognition and Segmentation”, Lecture Notes in Computer Science. Springer Berlin Heidelberg, pp. 1–15. Available at: https://doi.org/10.1007/11744023_1.
Vancouver
1. Shotton J, Winn J, Rother C, Criminisi A (2006) TextonBoost: Joint Appearance, Shape and Context Modeling for Multi-class Object Recognition and Segmentation. In: Lecture Notes in Computer Science. Springer Berlin Heidelberg, pp 1–15

BibTeX

@inbook{Shotton_2006, title={TextonBoost: Joint Appearance, Shape and Context Modeling for Multi-class Object Recognition and Segmentation}, ISBN={9783540338338}, ISSN={1611-3349}, url={http://dx.doi.org/10.1007/11744023_1}, DOI={10.1007/11744023_1}, booktitle={Computer Vision – ECCV 2006}, publisher={Springer Berlin Heidelberg}, author={Shotton, Jamie and Winn, John and Rother, Carsten and Criminisi, Antonio}, year={2006}, pages={1–15} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF