Conditional Random Fields as Recurrent Neural Networks

Shuai ZhengSadeep JayasumanaBernardino Romera-ParedesVibhav VineetZhizhong SuDalong DuChang HuangPhilip H. S. Torr

article2015ICCV2,609 citationsBest Demo Award

Integrates Conditional Random Fields into convolutional neural networks by formulating mean-field inference as recurrent layers, enabling end-to-end training for accurate semantic segmentation without disconnected post-processing.

Listen

Pixel-level image labelling tasks such as semantic segmentation require both accurate per-pixel predictions and consistent object boundaries. Traditional convolutional neural networks excel at feature learning for recognition but often produce coarse outputs with blurry edges and inconsistent labels because they lack explicit mechanisms for spatial smoothness and label agreement.

The article set out to create a unified deep network that integrates the feature-extraction power of CNNs with the boundary-refining capability of conditional random fields, allowing the entire system to be trained end-to-end rather than applying CRF inference only after CNN training.

The authors reformulated the iterative mean-field inference algorithm for dense CRFs with Gaussian pairwise potentials as a recurrent neural network layer. This CRF-RNN layer was inserted after a fully convolutional network (based on FCN-8s) so that unary potentials from the CNN are refined by the CRF component; gradients flow back through the RNN during training, jointly optimising both parts. Experiments used the Pascal VOC 2012 training set (augmented with extra annotations) and, in one variant, additional Microsoft COCO images, with evaluation on the VOC test set and the Pascal Context dataset.

The integrated model reached 74.7 percent mean intersection-over-union on the VOC 2012 test set, a new state-of-the-art at the time and roughly 2 percentage points above the best prior disconnected CRF approach. End-to-end training improved accuracy by 3.4 points over an otherwise identical network whose CRF parameters were fixed after initialisation. Using class-specific filter weights and a learned asymmetric label-compatibility function each contributed measurable gains. On the Pascal Context dataset the method also outperformed previous CNN-only and two-stage baselines.

These results show that allowing the CNN and CRF components to adapt to each other during training produces sharper boundaries and fewer spurious regions without sacrificing recognition accuracy. The improvement matters for applications such as autonomous driving or medical imaging where precise object delineation directly affects downstream decisions.

The authors recommend exploring whether more flexible recurrent architectures such as LSTMs could replace the mean-field iterations while preserving end-to-end trainability. They note that increasing the number of mean-field iterations beyond five during training can weaken gradients reaching the CNN, and that performance gains depend on careful initialisation of CRF parameters. The reported gains are consistent across two datasets, yet further validation on larger and more diverse imagery would strengthen confidence before deployment.

Cover for Conditional Random Fields as Recurrent Neural Networks

Abstract

Pixel-level labelling tasks, such as semantic segmentation, play a central role in image understanding. Recent approaches have attempted to harness the capabilities of deep learning techniques for image recognition to tackle pixel-level labelling tasks. One central issue in this methodology is the limited capacity of deep learning techniques to delineate visual objects. To solve this problem, we introduce a new form of convolutional neural network that combines the strengths of Convolutional Neural Networks (CNNs) and Conditional Random Fields (CRFs)-based probabilistic graphical modelling. To this end, we formulate mean-field approximate inference for the Conditional Random Fields with Gaussian pairwise potentials as Recurrent Neural Networks. This network, called CRF-RNN, is then plugged in as a part of a CNN to obtain a deep network that has desirable properties of both CNNs and CRFs. Importantly, our system fully integrates CRF modelling with CNNs, making it possible to train the whole deep network end-to-end with the usual back-propagation algorithm, avoiding offline post-processing methods for object delineation. We apply the proposed method to the problem of semantic image segmentation, obtaining top results on the challenging Pascal VOC 2012 segmentation benchmark.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Conditional Random Fields
  • 4. A Mean-field Iteration as a Stack of CNN Layers
  • 4.1. Initialization
  • 4.2. Message Passing
  • 4.3. Weighting Filter Outputs
  • 4.4. Compatibility Transform
  • 4.5. Adding Unary Potentials
  • 4.6. Normalization
  • 5. The End-to-end Trainable Network
  • 5.1. CRF as RNN
  • 5.2. Completing the Picture
  • 6. Implementation Details
  • 7. Experiments
  • Pascal VOC Datasets
  • Pascal Context Dataset
  • 7.1. Effect of Design Choices
  • 8. Conclusion
  • References

Knowls

  1. Knowl 1 — CRF-RNN Framework for Structured Semantic Segmentation

    model/method

    The CRF-RNN framework reformulates mean-field approximate inference in a fully connected (dense) Conditional Random Field (CRF) with Gaussian pairwise potentials as a Recurrent Neural Network (RNN). By expressing the iterative inference algorithm as a differentiable computation graph with recurrence, the CRF graphical model is integrated directly on top of a Convolutional Neural Network (CNN) producing unary potentials (such as an FCN-8s architecture).

    This unified model allows the parameters of both the CNN feature extractor and the CRF (filter kernel weights and label compatibility matrices) to be trained jointly end-to-end via standard backpropagation through time (BPTT) and stochastic gradient descent (SGD). This avoids the sub-optimality of treating dense CRF inference as an offline, disconnected post-processing step.

  2. Knowl 2 — Mean-Field CRF Inference as a Stack of Differentiable CNN Layers

    model/method

    A single iteration of mean-field approximate inference for dense CRFs with Gaussian edge potentials is decomposed into five consecutive, fully differentiable neural network operations:

    1. Message Passing: High-dimensional Gaussian filtering is applied to the estimated marginal probability distributions QQ. For an image with NN pixels, MM Gaussian kernels (specifically, a spatial kernel and a bilateral appearance kernel) compute filtered responses Q~i(m)(l)=jik(m)(fi,fj)Qj(l)\tilde{Q}^{(m)}_i(l) = \sum_{j \neq i} k^{(m)}(f_i, f_j) Q_j(l) across all pixels i,ji, j and labels ll. To make this operation tractable across the entire image in O(N)O(N) time, a permutohedral lattice is employed for both the forward pass and the backward pass (by reversing the order of separable filters during lattice blurring).

    2. Weighting Filter Outputs: A weighted linear combination of the MM filter outputs is computed for each label ll: Qˇi(l)=m=1Mwl(m)Q~i(m)(l)\check{Q}_i(l) = \sum_{m=1}^M w^{(m)}_l \tilde{Q}^{(m)}_i(l). This operation is implemented as a 1×11 \times 1 convolution with MM input channels and 11 output channel per class label, enabling class-specific weights wl(m)w^{(m)}_l to be learned.

    3. Compatibility Transform: Label interactions are modeled by Q^i(l)=lLμ(l,l)Qˇi(l)\hat{Q}_i(l) = \sum_{l' \in \mathcal{L}} \mu(l, l') \check{Q}_i(l'), where μ(l,l)\mu(l, l') is a learnable label compatibility function. This is implemented as a 1×11 \times 1 convolution layer with LL input and LL output channels, generalized to permit asymmetric label penalties μ(l,l)μ(l,l)\mu(l, l') \neq \mu(l', l).

    4. Adding Unary Potentials: The compatibility-transformed outputs are subtracted element-wise from the negative unary potentials Ui(l)=ψu(Xi=l)U_i(l) = -\psi_u(X_i = l) provided by the preceding CNN: Q˘i(l)=Ui(l)Q^i(l)\breve{Q}_i(l) = U_i(l) - \hat{Q}_i(l). Error differentials are propagated directly with inverted signs during backpropagation.

    5. Normalization: The updated marginal distribution is obtained by computing a pixel-wise softmax across all label channels: Qi(l)=1Ziexp(Q˘i(l))Q_i(l) = \frac{1}{Z_i} \exp(\breve{Q}_i(l)), where Zi=lLexp(Q˘i(l))Z_i = \sum_{l' \in \mathcal{L}} \exp(\breve{Q}_i(l')).

  3. Knowl 3 — Dense CRF Mean-Field Inference as CNN Operations

    algorithm

    The algorithm executes iterative mean-field inference on dense CRFs using operations structured as differentiable CNN layers:

    Algorithm: Dense CRF Mean-Field Inference via CNN Operations
    Input: Unary potential map URN×LU \in \mathbb{R}^{N \times L} where Ui(l)=ψu(Xi=l)U_i(l) = -\psi_u(X_i = l) for pixel ii and label lL={l1,,lL}l \in \mathcal{L} = \{l_1, \dots, l_L\}; image feature vectors {fi}i=1N\{f_i\}_{i=1}^N; number of iterations TT; Gaussian kernels {k(m)}m=1M\{k^{(m)}\}_{m=1}^M; learnable weights wl(m)w^{(m)}_l; learnable compatibility matrix μRL×L\mu \in \mathbb{R}^{L \times L}.
    Output: Marginal probability distributions QRN×LQ \in \mathbb{R}^{N \times L}.
    for each pixel i{1,,N}i \in \{1, \dots, N\} and label lLl \in \mathcal{L} do
        Qi(l)exp(Ui(l))lLexp(Ui(l))Q_i(l) \leftarrow \frac{\exp(U_i(l))}{\sum_{l' \in \mathcal{L}} \exp(U_i(l'))}
    end for
    for t=1t = 1 to TT do
        for each kernel m{1,,M}m \in \{1, \dots, M\} and label lLl \in \mathcal{L} do
            Q~i(m)(l)jik(m)(fi,fj)Qj(l)\tilde{Q}_i^{(m)}(l) \leftarrow \sum_{j \neq i} k^{(m)}(f_i, f_j) Q_j(l) via permutohedral lattice
        end for
        for each label lLl \in \mathcal{L} do
            Qˇi(l)m=1Mwl(m)Q~i(m)(l)\check{Q}_i(l) \leftarrow \sum_{m=1}^M w^{(m)}_l \tilde{Q}_i^{(m)}(l)
        end for
        for each label lLl \in \mathcal{L} do
            Q^i(l)lLμ(l,l)Qˇi(l)\hat{Q}_i(l) \leftarrow \sum_{l' \in \mathcal{L}} \mu(l, l') \check{Q}_i(l')
        end for
        for each label lLl \in \mathcal{L} do
            Q˘i(l)Ui(l)Q^i(l)\breve{Q}_i(l) \leftarrow U_i(l) - \hat{Q}_i(l)
        end for
        for each label lLl \in \mathcal{L} do
            Qi(l)exp(Q˘i(l))lLexp(Q˘i(l))Q_i(l) \leftarrow \frac{\exp(\breve{Q}_i(l))}{\sum_{l' \in \mathcal{L}} \exp(\breve{Q}_i(l'))}
        end for
    end for
    return QQ
  4. Knowl 4 — Recurrent State Equations of CRF-RNN

    equation

    Let II be the input image, UU be the unary potential tensor generated by the front-end CNN, and fθ(U,Qin,I)f_\theta(U, Q_{\text{in}}, I) denote the transformation performed by one mean-field iteration parameterized by θ={wl(m),μ(l,l)}\theta = \{w^{(m)}_l, \mu(l, l')\}. For a total number of mean-field iterations TT, the recurrent dynamics of the CRF-RNN are governed by:

    H1(t)={softmax(U),t=0H2(t1),0<tTH_1(t) = \begin{cases} \operatorname{softmax}(U), & t = 0 \\ H_2(t - 1), & 0 < t \le T \end{cases} H2(t)=fθ(U,H1(t),I),0tTH_2(t) = f_\theta(U, H_1(t), I), \quad 0 \le t \le T Y(t)={0,0t<TH2(t),t=TY(t) = \begin{cases} 0, & 0 \le t < T \\ H_2(t), & t = T \end{cases}

    Here H1(t)H_1(t) represents the input distribution estimate fed into iteration tt, H2(t)H_2(t) is the output distribution estimate after iteration tt, and Y(t)Y(t) is the final output gated to emit results only at time step t=Tt = T.

  5. Knowl 5 — End-to-End CNN and CRF-RNN Integration and Training Setup

    experimental setup

    The complete network consists of an FCN-8s network (derived from VGG-16) providing unary potentials, directly followed by the CRF-RNN block and terminated with a softmax loss layer on pixel ground truth.

    Key implementation and optimization parameters include:

    • Initialization: CNN layers are initialized with pre-trained FCN-8s weights. CRF compatibility transform weights are initialized using the Potts model (1-1 on diagonals, 00 off-diagonals). Kernel weights and filter bandwidths are initialized via cross-validation.
    • Optimization: Trained via Stochastic Gradient Descent (SGD) using whole-image mini-batches (batch size 1) with momentum fixed at 0.990.99 and a learning rate of 101310^{-13}.
    • Iterations: Mean-field iterations are set to T=5T = 5 during training to prevent vanishing/exploding gradients and reduce compute time, and increased to T=10T = 10 during inference.
    • Loss Function: Standard pixel-wise log-likelihood / softmax loss.
  6. Knowl 6 — Comparative Performance of End-to-End CRF-RNN vs. Disconnected CRF Inference

    data/table

    Evaluating different training strategies on the reduced Pascal VOC 2012 validation set (346 images without training overlap) demonstrates the quantitative advantage of end-to-end joint training of CNN and CRF-RNN compared to using CRF as a standalone, disconnected post-processing step on top of FCN-8s outputs:

    Method Without COCO With COCO
    Plain FCN-8s 61.3 68.3
    FCN-8s and CRF disconnected 63.7 69.5
    End-to-end training of CRF-RNN 69.6 72.9

    All numbers report mean Intersection over Union (IU, %). End-to-end training improves over disconnected CRF post-processing by +5.9% IU without COCO pretraining and by +3.4% IU with COCO pretraining.

  7. Knowl 7 — Semantic Segmentation Benchmark Results on Pascal VOC 2010-2012

    data/table

    Performance of CRF-RNN compared to competing state-of-the-art segmentation methods on the Pascal VOC 2010, 2011, and 2012 test benchmarks:

    Method VOC 2010 test VOC 2011 test VOC 2012 test
    Without COCO training data:
    BerkeleyRC n/a 39.1 n/a
    O2PCPMC 49.6 48.8 47.8
    Divmbest n/a n/a 48.1
    NUS-UDS n/a n/a 50.0
    SDS n/a n/a 51.6
    MSRA-CFM n/a n/a 61.8
    FCN-8s n/a 62.7 62.2
    Hypercolumn n/a n/a 62.6
    Zoomout 64.4 64.1 64.4
    Context-Deep-CNN-CRF n/a n/a 70.7
    DeepLab-MSc n/a n/a 71.6
    CRF-RNN (w/o COCO) 73.6 72.4 72.0
    With COCO pre-training:
    BoxSup n/a n/a 71.0
    DeepLab n/a n/a 72.7
    CRF-RNN (with COCO) 75.7 75.0 74.7

    Metric reported is mean Intersection over Union (IU, %). CRF-RNN achieves state-of-the-art accuracy across all three VOC test sets (72.0% without COCO, 74.7% with COCO on VOC 2012).

  8. Knowl 8 — Semantic Segmentation Evaluation on the Pascal Context Dataset

    data/table

    Performance comparison on the 59-class Pascal Context dataset validation partition:

    Method O2PO_2P CFM FCN-8s CRF-RNN
    Mean IU (%) 18.1 34.4 37.78 39.28

    CRF-RNN improves upon FCN-8s by 1.50% mean IU on this 59-class scene parsing benchmark.

  9. Knowl 9 — Ablation Analysis of CRF-RNN Architectural Components and Parameters

    empirical result

    Empirical ablation on the Pascal VOC 2012 validation set reveals the impact of individual design choices:

    • Class-Dependent Kernel Weights: Learning independent Gaussian filter weights wl(m)w^{(m)}_l for each class ll yields a +1.8 percentage point improvement over class-agnostic weights.
    • Asymmetric Label Compatibility: Allowing an asymmetric compatibility matrix μ(l,l)μ(l,l)\mu(l, l') \neq \mu(l', l) provides an additional +0.9 percentage point boost.
    • End-to-End Joint Training vs. Partial Tuning: End-to-end training of the combined network after initialization gains +3.4 percentage points, whereas freezing the CNN and fine-tuning only the CRF parameters improves accuracy by only +1.0 percentage point.
    • Recurrent Weight Sharing vs. Independent Layers: Unrolling 5 iterations with independent parameters per iteration produces an IU of 70.9%, underperforming the recurrent weight-sharing formulation and indicating the importance of strict recurrence.
    • Iteration Count TT: Increasing test iterations from T=5T = 5 to T=10T = 10 improves accuracy by +0.2 percentage points. Increasing training iterations to T=10T = 10 reduces accuracy by 0.7 percentage points due to vanishing gradients propagating through the RNN loop to the unary CNN.
    • Softmax vs. ReLU Normalization: Replacing exponential/softmax normalization with a ReLU followed by normalization degraded performance by 1.0% IU.
  10. Knowl 10 — Asymmetric Label Compatibility Formulation in Dense CRFs

    model/method

    In standard dense CRF formulations using the Potts model, label compatibility is binary and symmetric: μ(l,l)=[ll]\mu(l, l') = [l \neq l'], applying an identical penalty to all distinct label co-occurrences.

    CRF-RNN replaces this with a fully learnable, unconstrained L×LL \times L weight matrix implemented as a 1×11 \times 1 convolution across class channels, where μ(l,l)μ(l,l)\mu(l, l') \neq \mu(l', l) is permitted. This enables the network to learn asymmetric co-occurrence priors from training data (e.g., assigning lower penalties to mutually consistent pairs such as [Motorbike, Person] or [Dining table, Chair] compared to unrelated pairs like [Sky, Bicycle]).

Coverage note — None was omitted; all primary architectural components, algorithmic steps, equations, training setups, ablation studies, and experimental results on Pascal VOC and Pascal Context are covered.

References

  1. 1.A. Adams, J. Baek, and M. A. Davis. Fast high-dimensional filtering using the permutohedral lattice. Computer Graphics Forum, 29(2):753–762, 2010.
  2. 2.P. Arbelaez, B. Hariharan, C. Gu, S. Gupta, L. Bourdev, and J. Malik. Semantic segmentation using regions and parts. In IEEE CVPR, 2012.
  3. 3.P. Arbelaez, M. Maire, C. Fowlkes, and J. Malik. Contour detection and hierarchical image segmentation. IEEE TPAMI, 33(5):898–916, 2011.
  4. 4.A. Barbu. Training an active random field for real-time image denoising. IEEE TIP, 18(11):2451–2462, 2009.
  5. 5.S. Bell, P. Upchurch, N. Snavely, and K. Bala. Material recognition in the wild with the materials in context database. In IEEE CVPR, 2015.
  6. 6.Y. Bengio, Y. LeCun, and D. Henderson. Globally trained handwritten word recognizer using spatial representation, convolutional neural networks, and hidden markov models. In NIPS, pages 937–937, 1994.
  7. 7.Y. Bengio, P. Simard, and P. Frasconi. Learning long-term dependencies with gradient descent is difficult. IEEE Transactions on Neural Networks, 1994.
  8. 8.J. Carreira, R. Caseiro, J. Batista, and C. Sminchisescu. Free-form region description with second-order pooling. IEEE TPAMI, 2014.
  9. 9.K. Chatfield, K. Simonyan, A. Vedaldi, and A. Zisserman. Return of the devil in the details: Delving deep into convolutional nets. In BMVC, 2014.
  10. 10.L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille. Semantic image segmentation with deep convolutional nets and fully connected crfs. In ICLR, 2015.
  11. 11.L.-C. Chen, A. G. Schwing, A. L. Yuille, and R. Urtasun. Learning deep structured models. In ICLRW, 2015.
  12. 12.J. Dai, K. He, and J. Sun. Boxsup: Exploiting bounding boxes to supervise convolutional networks for semantic segmentation. In arXiv:1503.01640, 2015.
  13. 13.J. Dai, K. He, and J. Sun. Convolutional feature masking for joint object and stuff segmentation. In IEEE CVPR, 2015.
  14. 14.T.-M.-T. Do and T. Artieres. Neural conditional random fields. In NIPS, 2010.
  15. 15.J. Domke. Learning graphical model parameters with approximate marginal inference. IEEE TPAMI, 35(10):2454–2467, 2013.
  16. 16.J. Dong, Q. Chen, S. Yan, and A. Yuille. Towards unified object detection and semantic segmentation. In ECCV, 2014.
  17. 17.D. Eigen, C. Puhrsch, and R. Fergus. Depth map prediction from a single image using a multi-scale deep network. In NIPS, 2014.
  18. 18.M. Everingham, S. M. A. Eslami, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman. The pascal visual object classes challenge: A retrospective. IJCV, 111(1):98–136, 2015.
  19. 19.C. Farabet, C. Couprie, L. Najman, and Y. LeCun. Learning hierarchical features for scene labeling. IEEE TPAMI, 2013.
  20. 20.R. Girshick, J. Donahue, T. Darrell, and J. Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In IEEE CVPR, 2014.
  21. 21.R. Girshick, F. Iandola, T. Darrell, and J. Malik. Deformable part models are convolutional neural networks. In CVPR, 2015.
  22. 22.B. Hariharan, P. Arbelaez, L. D. Bourdev, S. Maji, and J. Malik. Semantic contours from inverse detectors. In IEEE ICCV, 2011.
  23. 23.B. Hariharan, P. Arbeláez, R. Girshick, and J. Malik. Simultaneous detection and segmentation. In ECCV, 2014.
  24. 24.B. Hariharan, P. Arbelaez, R. Girshick, and J. Malik. Hypercolumns for object segmentation and fine-grained localization. In IEEE CVPR, 2015.
  25. 25.M. Jaderberg, K. Simonyan, A. Vedaldi, and A. Zisserman. Deep structured output learning for unconstrained text recognition. In ICLR, 2015.
  26. 26.V. Jain, J. F. Murray, F. Roth, S. C. Turaga, V. P. Zhigulin, K. L. Briggman, M. Helmstaedter, W. Denk, and H. S. Seung. Supervised learning of image restoration with convolutional networks. In IEEE ICCV, 2007.
  27. 27.Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell. Caffe: Convolutional architecture for fast feature embedding. In ACM Multimedia, pages 675–678, 2014.
  28. 28.M. Kiefel and P. V. Gehler. Human pose estmation with fields of parts. In ECCV, 2014.
  29. 29.P. Krähenbühl and V. Koltun. Efficient inference in fully connected crfs with gaussian edge potentials. In NIPS, 2011.
  30. 30.P. Krähenbühl and V. Koltun. Parameter learning and convergent inference for dense random fields. In ICML, 2013.
  31. 31.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In NIPS, 2012.
  32. 32.L. Ladicky, C. Russell, P. Kohli, and P. H. Torr. Associative hierarchical crfs for object class image segmentation. In IEEE ICCV, 2009.
  33. 33.J. D. Lafferty, A. McCallum, and F. C. N. Pereira. Conditional random fields: Probabilistic models for segmenting and labeling sequence data. In ICML, 2001.
  34. 34.Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998.
  35. 35.G. Lin, C. Shen, I. Reid, and A. van dan Hengel. Efficient piecewise training of deep structured models for semantic segmentation. In arXiv:1504.01013, 2015.
  36. 36.T.-Y. Lin, M. Maire, S. Belongie, L. Bourdev, R. Girshick, J. Hays, P. Perona, D. Ramanan, C. L. Zitnick, and P. Dollar. Microsoft coco: Common objects in context. In arXiv:1405.0312, 2014.
  37. 37.J. Long, E. Shelhamer, and T. Darrell. Fully convolutional networks for semantic segmentation. In IEEE CVPR, 2015.
  38. 38.M. Mostajabi, P. Yadollahpour, and G. Shakhnarovich. Feedforward semantic segmentation with zoom-out features. In IEEE CVPR, 2015.
  39. 39.R. Mottaghi, X. Chen, X. Liu, N.-G. Cho, S.-W. Lee, S. Fidler, R. Urtasun, and A. Yuille. The role of context for object detection and semantic segmentation in the wild. In IEEE CVPR, 2014.
  40. 40.M. C. Mozer. Backpropagation. chapter A Focused Backpropagation Algorithm for Temporal Pattern Recognition. L. Erlbaum Associates Inc., 1995.
  41. 41.G. Papandreou, L.-C. Chen, K. Murphy, and A. L. Yuille. Weakly- and semi-supervised learning of a dcnn for semantic image segmentation. In arXiv:1502.02734, 2015.
  42. 42.S. Paris and F. Durand. A fast approximation of the bilateral filter using a signal processing approach. IJCV, 81(1):24–52, 2013.
  43. 43.R. Pascanu, C. Gulcehre, K. Cho, and Y. Bengio. On the difficulty of training recurrent neural networks. In ICML, 2013.
  44. 44.G. S. Payman Yadollahpour, Dhruv Batra. Discriminative re-ranking of diverse segmentations. In IEEE CVPR, 2013.
  45. 45.J. Peng, L. Bo, and J. Xu. Conditional neural fields. In NIPS, 2009.
  46. 46.P. H. O. Pinheiro and R. Collobert. Recurrent convolutional neural networks for scene labeling. In ICML, 2014.
  47. 47.S. Ross, D. Munoz, M. Hebert, and J. A. Bagnell. Learning message-passing inference machines for structured prediction. In IEEE CVPR, 2011.
  48. 48.D. E. Rumelhart, G. E. Hinton, and R. J. Williams. Neurocomputing: Foundations of research. chapter Learning Internal Representations by Error Propagation. MIT Press, 1988.
  49. 49.A. G. Schwing and R. Urtasun. Fully connected deep structured networks. In arXiv:1503.02351, 2015.
  50. 50.J. Shotton, A. Fitzgibbon, M. Cook, T. Sharp, M. Finocchio, R. Moore, A. Kipman, and A. Blake. Real-time human pose recognition in parts from single depth images. In IEEE CVPR, 2011.
  51. 51.J. Shotton, M. Johnson, and R. Cipolla. Semantic texton forests for image categorization and segmentation. In IEEE CVPR, 2008.
  52. 52.J. Shotton, J. Winn, C. Rother, and A. Criminisi. Textonboost for image understanding: Multi-class object recognition and segmentation by jointly modeling texture, layout, and context. IJCV, 81(1):2–23, 2009.
  53. 53.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. In arXiv:1409.1556, 2014.
  54. 54.V. Stoyanov, A. Ropson, and J. Eisner. Empirical risk minimization of graphical model parameters given approximate inference, decoding, and model structure. In AISTATS, 2011.
  55. 55.S. C. Tatikonda and M. I. Jordan. Loopy belief propagation and gibbs measures. In Proceedings of the Eighteenth Conference on Uncertainty in Artificial Intelligence, 2002.
  56. 56.C. Tomasi and R. Manduchi. Bilateral filtering for gray and color images. In IEEE CVPR, 1998.
  57. 57.J. J. Tompson, A. Jain, Y. LeCun, and C. Bregler. Joint training of a convolutional network and a graphical model for human pose estimation. In NIPS, 2014.
  58. 58.Z. Tu. Auto-context and its application to high-level vision tasks. In IEEE CVPR, 2008.
  59. 59.Z. Tu, X. Chen, A. L. Yuille, and S.-C. Zhu. Image parsing: Unifying segmentation, detection, and recognition. IJCV, 63(2):113–140, 2005.
  60. 60.K. Yao, B. Peng, G. Zweig, D. Yu, X. Li, and F. Gao. Recurrent conditional random field for language understanding. In ICASSP, 2014.
  61. 61.Y. Zhang and T. Chen. Efficient inference for fullyconnected crfs with stationarity. In CVPR, 2012.

Citation

MLA
Zheng, S., et al. “Conditional Random Fields as Recurrent Neural Networks”. 2015 IEEE International Conference on Computer Vision (ICCV), 2015, pp. 1529–37, https://doi.org/10.1109/ICCV.2015.179.
APA
Zheng, S., Jayasumana, S., Romera-Paredes, B., Vineet, V., Su, Z., Du, D., Huang, C., & Torr, P. H. S. (2015). Conditional Random Fields as Recurrent Neural Networks. 2015 IEEE International Conference on Computer Vision (ICCV), 1529–1537. https://doi.org/10.1109/ICCV.2015.179
Chicago
Zheng, S., S. Jayasumana, B. Romera-Paredes, et al. 2015. “Conditional Random Fields as Recurrent Neural Networks”. 2015 IEEE International Conference on Computer Vision (ICCV), 1529–37. https://doi.org/10.1109/ICCV.2015.179.
Harvard
Zheng, S. et al. (2015) “Conditional Random Fields as Recurrent Neural Networks”, 2015 IEEE International Conference on Computer Vision (ICCV). IEEE, pp. 1529–1537. Available at: https://doi.org/10.1109/ICCV.2015.179.
Vancouver
1. Zheng S, Jayasumana S, Romera-Paredes B, Vineet V, Su Z, Du D, Huang C, Torr PHS (2015) Conditional Random Fields as Recurrent Neural Networks. In: 2015 IEEE International Conference on Computer Vision (ICCV). IEEE, pp 1529–1537

BibTeX

@inproceedings{Zheng_2015, title={Conditional Random Fields as Recurrent Neural Networks}, url={http://dx.doi.org/10.1109/ICCV.2015.179}, DOI={10.1109/iccv.2015.179}, booktitle={2015 IEEE International Conference on Computer Vision (ICCV)}, publisher={IEEE}, author={Zheng, Shuai and Jayasumana, Sadeep and Romera-Paredes, Bernardino and Vineet, Vibhav and Su, Zhizhong and Du, Dalong and Huang, Chang and Torr, Philip H. S.}, year={2015}, month=Dec, pages={1529–1537} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: IEEE