PIDNet: A Real-time Semantic Segmentation Network Inspired by PID Controllers

Jiacong XuZixiang XiongShankar P. Bhattacharyya

article2023CVPR664 citations

Proposes a three-branch real-time semantic segmentation architecture inspired by PID controllers that introduces a dedicated boundary branch to prevent contextual feature overshoot, achieving state-of-the-art speed-accuracy trade-offs on Cityscapes and CamVid.

Listen

Real-time visual scene parsing, which assigns an object category to every pixel in an image, is critical for high-stakes technologies like autonomous driving and robotic surgery. While existing two-branch deep learning models deliver fast processing speeds, they often suffer from an "overshoot" problem where surrounding contextual information overwhelms fine spatial details. This flaw causes boundary errors and causes the system to miss small objects, presenting safety and reliability risks in live applications.

The article introduces and evaluates PIDNet, a novel three-branch network architecture inspired by classical Proportional-Integral-Derivative (PID) control theory. The primary objective is to demonstrate that treating image parsing through a PID control framework mitigates detail loss and establishes a superior balance between inference speed and segmentation accuracy.

To achieve this, the authors designed a three-branch architecture: a Proportional branch for high-resolution details, an Integral branch for broad context, and an auxiliary Derivative branch dedicated to detecting object boundaries. Special fusion modules use boundary attention to prevent context from overpowering fine details, while a parallel pooling module accelerates context gathering. The system was trained and evaluated on standard image benchmarks—including Cityscapes, CamVid, and PASCAL Context—using standard graphics hardware to measure accuracy (mean Intersection-over-Union, or mIOU) and speed (frames per second, or FPS).

The evaluation produced four key findings. First, PIDNet established the highest accuracy among real-time models without using external hardware acceleration tools, with the large variant (PIDNet-L) reaching 80.6% mIOU at 31.1 FPS on the Cityscapes test set. Second, the compact model (PIDNet-S) demonstrated exceptional real-time efficiency, running at 93.2 FPS with 78.6% accuracy on Cityscapes, and 153.7 FPS with 80.1% accuracy on CamVid. Third, the model outperformed prior leading architectures; on CamVid, PIDNet-S improved accuracy by 1.5 percentage points over DDRNet-23-S with only about 1 millisecond of added latency. Finally, ablation tests confirmed that the derivative boundary branch and boundary-guided fusion modules were directly responsible for these performance gains, significantly sharpening edge delineation and improving small object detection.

These results demonstrate that control theory principles can effectively resolve core trade-offs in computer vision. For operational deployments, PIDNet lowers latency and computational overhead without sacrificing the visual fidelity necessary to reliably detect boundaries and obstacles. This balance reduces hardware costs while maintaining the safety standards required for autonomous navigation.

Decision-makers building real-time vision pipelines should consider adopting PIDNet-based models, choosing between the high-accuracy PIDNet-L and the high-speed PIDNet-S or PIDNet-S-Simple depending on latency constraints. As next steps, engineering teams should conduct pilot testing on target embedded hardware. Stakeholders should note that PIDNet relies on precise boundary annotations during training for optimal results, meaning data curation pipelines may require higher-quality edge labeling to achieve these gains in custom applications.

arXiv: 2206.02066
Cover for PIDNet: A Real-time Semantic Segmentation Network Inspired by PID Controllers

Abstract

Two-branch network architecture has shown its efficiency and effectiveness in real-time semantic segmentation tasks. However, direct fusion of high-resolution details and low-frequency context has the drawback of detailed features being easily overwhelmed by surrounding contextual information. This overshoot phenomenon limits the improvement of the segmentation accuracy of existing two-branch models. In this paper, we make a connection between Convolutional Neural Networks (CNN) and Proportional-Integral-Derivative (PID) controllers and reveal that a two-branch network is equivalent to a Proportional-Integral (PI) controller, which inherently suffers from similar overshoot issues. To alleviate this problem, we propose a novel three-branch network architecture: PIDNet, which contains three branches to parse detailed, context and boundary information, respectively, and employs boundary attention to guide the fusion of detailed and context branches. Our family of PIDNets achieve the best trade-off between inference speed and accuracy and their accuracy surpasses all the existing models with similar inference speed on the Cityscapes and CamVid datasets. Specifically, PIDNet-S achieves 78.6% mIOU with inference speed of 93.2 FPS on Cityscapes and 80.1% mIOU with speed of 153.7 FPS on CamVid.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. High-accuracy Semantic Segmentation
  • 2.2. Real-time Semantic Segmentation
  • 3. Method
  • 3.1. PIDNet: A Novel Three-branch Network
  • 3.2. Pag: Learning High-level Semantics Selectively
  • 3.3. PAPPM: Fast Aggregation of Contexts
  • 3.4. Bag: Balancing the Details and Contexts
  • 4. Experiment
  • 4.1. Datasets
  • 4.2. Implementation Details
  • 4.3. Ablation Study
  • 4.4. Comparison
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — PIDNet Three-Branch Architecture

    model/method

    PIDNet (Proportional-Integral-Derivative Network) is a three-branch real-time semantic segmentation neural network designed to resolve the overshoot phenomenon caused by the direct fusion of high-resolution details and low-frequency contextual representations in two-branch architectures.

    The network comprises three complementary branches operating at distinct depths and resolutions:

    • Proportional (P) Branch: A moderate-depth branch operating on high-resolution feature maps (1/81/8 of the input spatial resolution) to preserve fine spatial details and delineate object boundaries.
    • Integral (I) Branch: A deep branch utilizing strided residual blocks downsampled successively to 1/16,1/32,1/16, 1/32, and 1/641/64 resolutions, capped with a context aggregation block (PAPPM or DAPPM) to aggregate local and global semantic context and capture long-range dependencies.
    • Derivative (D) Branch: A shallow auxiliary branch operating at 1/81/8 spatial resolution to extract high-frequency spatial differences and predict boundary regions.

    The Integral branch transfers semantic information into the Proportional branch via Pixel-attention-guided fusion modules (Pag) and into the Derivative branch via element-wise addition. The Proportional and Integral feature maps are finally fused using a Boundary-attention-guided fusion module (Bag or Light-Bag) modulated by the spatial boundary probability generated by the Derivative branch.

    The architecture is instantiated in three scale configurations—PIDNet-S, PIDNet-M, and PIDNet-L—by adjusting depth and channel widths, as well as an ultra-fast variant PIDNet-S-Simple that strips Pag and Bag modules for minimum latency.

  2. Knowl 2 — Analogy Between Two-Branch Networks and PI Controllers

    theoretical result

    In two-branch real-time semantic segmentation networks, the detail branch preserves high-resolution local features while the context branch downsamples feature maps to extract global context. In a multi-layer 1D linear approximation without batch normalization or activation functions, the effective convolution kernel coefficients decay exponentially with layer depth. As a result, the detail branch concentrates over 70%70\% of its effective kernel weight within the immediate spatial neighborhood ({I[i−1],I[i],I[i+1]}\{I[i-1], I[i], I[i+1]\}), whereas the context branch spreads its weight across larger spatial extents with less than 26%26\% falling in the local neighborhood.

    In the frequency domain (z=ejωz = e^{j\omega}), the zz-transform of a PID controller is expressed as: C(z)=kp+ki(1−e−jω)−1+kd(1−e−jω)C(z) = k_p + k_i(1 - e^{-j\omega})^{-1} + k_d(1 - e^{-j\omega}) where ω\omega denotes spatial frequency, kp≥0k_p \ge 0 is the proportional gain, ki≥0k_i \ge 0 is the integral gain, and kd≥0k_d \ge 0 is the derivative gain. As ω\omega increases, the proportional component acts as an allpass filter, the integral component acts as a lowpass filter with decaying gain, and the derivative component acts as a highpass filter with increasing gain.

    A two-branch network is mathematically analogous to a Proportional-Integral (PI) controller: the context branch functions as an averaging lowpass filter that aggregates semantics. Because PI controllers accumulate low-frequency inputs and cannot react promptly to sudden signal changes, they suffer from overshoot. In semantic segmentation, this manifests as spatial overshoot where object boundaries are corroded by adjacent low-frequency context and small objects are overwhelmed by large surrounding regions. Introducing an auxiliary derivative branch that extracts high-frequency semantic variations (boundary detection) acts as a highpass damper that suppresses this overshoot.

  3. Knowl 3 — Pixel-Attention-Guided Fusion Module

    model/method

    The Pixel-attention-guided fusion module (Pag) enables the high-resolution Proportional (P) branch to selectively incorporate high-level semantic context from the deep Integral (I) branch without having local spatial details overwhelmed.

    Let v⃗p∈RC\vec{v}_p \in \mathbb{R}^C and v⃗i∈RC\vec{v}_i \in \mathbb{R}^C denote feature vectors at corresponding spatial pixel locations in the P branch and I branch, respectively. Let fp(⋅)f_p(\cdot) and fi(⋅)f_i(\cdot) be linear projections implemented with 1×11 \times 1 convolutions followed by batch normalization. The pixel attention score σ∈(0,1)\sigma \in (0, 1) is computed as: σ=Sigmoid(fp(v⃗p)⋅fi(v⃗i))\sigma = \text{Sigmoid}\left( f_p(\vec{v}_p) \cdot f_i(\vec{v}_i) \right) where ⋅\cdot denotes vector inner product. The output vector OutPag∈RC\text{Out}_{Pag} \in \mathbb{R}^C of the module is: OutPag=σv⃗i+(1−σ)v⃗p\text{Out}_{Pag} = \sigma \vec{v}_i + (1 - \sigma) \vec{v}_p When σ\sigma is close to 1 (indicating high semantic agreement that both features represent the same object), the output prioritizes rich contextual features from the I branch; when σ\sigma is close to 0, it retains local detail from the P branch.

  4. Knowl 4 — Boundary-Attention-Guided Fusion Modules (Bag and Light-Bag)

    model/method

    The Boundary-attention-guided fusion module (Bag) and its lightweight counterpart (Light-Bag) fuse spatial details from the Proportional (P) branch and context representations from the Integral (I) branch using the boundary predictions from the Derivative (D) branch.

    Let v⃗p∈RC\vec{v}_p \in \mathbb{R}^C, v⃗i∈RC\vec{v}_i \in \mathbb{R}^C, and v⃗d∈R\vec{v}_d \in \mathbb{R} represent feature vectors at the same spatial position from the P, I, and D branches, respectively. The boundary confidence score σ∈(0,1)\sigma \in (0, 1) is: σ=Sigmoid(v⃗d)\sigma = \text{Sigmoid}(\vec{v}_d)

    In the standard Bag module: Outbag=fout((1−σ)⊗v⃗i+σ⊗v⃗p)\text{Out}_{bag} = f_{out}\left( (1 - \sigma) \otimes \vec{v}_i + \sigma \otimes \vec{v}_p \right) where ⊗\otimes denotes channel-wise broadcast multiplication, and foutf_{out} is a sequence of 3×33 \times 3 convolution, batch normalization, and ReLU.

    In the Light-Bag module (designed for lightweight models such as PIDNet-S to minimize latency): Outlight=fp((1−σ)⊗v⃗i+v⃗p)+fi(σ⊗v⃗p+v⃗i)\text{Out}_{light} = f_p\left( (1 - \sigma) \otimes \vec{v}_i + \vec{v}_p \right) + f_i\left( \sigma \otimes \vec{v}_p + \vec{v}_i \right) where fpf_p and fif_i consist of 1×11 \times 1 convolutions, batch normalizations, and ReLUs.

    When σ>0.5\sigma > 0.5 (indicating high boundary likelihood), the module emphasizes detailed features from the P branch; when σ≤0.5\sigma \le 0.5, context features from the I branch dominate the representation.

  5. Knowl 5 — Parallel Aggregation Pyramid Pooling Module

    model/method

    The Parallel Aggregation Pyramid Pooling Module (PAPPM) is an efficient context harvesting block deployed in PIDNet-S and PIDNet-M to construct multi-scale scene representations while enabling parallel computation on hardware.

    PAPPM addresses the sequential execution latency and heavy channel dimensionality of Deep Aggregation PPM (DAPPM):

    1. The input feature map is pooled in parallel across multiple scales using average pooling: global pooling, 17×1717 \times 17 pooling with stride 8, 9×99 \times 9 pooling with stride 4, and 5×55 \times 5 pooling with stride 2.
    2. Each pooled output is mapped by a 1×11 \times 1 convolution, upsampled bilinearly back to the input spatial dimensions, and merged hierarchically across scales via element-wise addition.
    3. The outputs from all branches and a 1×11 \times 1 convolutional shortcut of the input are concatenated and projected through a final 1×11 \times 1 convolution.

    PAPPM reduces the internal channel dimension per pooling scale from 128 to 96 and executes all branches concurrently, delivering a +9.5+9.5 FPS speedup over DAPPM in PIDNet-S while achieving identical segmentation accuracy (78.8%78.8\% mIoU on Cityscapes validation).

  6. Knowl 6 — Multi-Loss Training Objective for PIDNet

    equation

    PIDNet is trained end-to-end using a weighted composite loss combining four distinct supervision objectives: Ltotal=λ0l0+λ1l1+λ2l2+λ3l3\mathcal{L}_{\text{total}} = \lambda_0 l_0 + \lambda_1 l_1 + \lambda_2 l_2 + \lambda_3 l_3 where:

    • l0l_0 is an auxiliary cross-entropy loss computed from a semantic head attached to the output of the first Pag module.
    • l1l_1 is a weighted binary cross-entropy loss applied to the boundary head of the Derivative (D) branch to supervise coarse object boundary detection.
    • l2l_2 is the standard cross-entropy loss applied to the final semantic segmentation output.
    • l3l_3 is the boundary-awareness semantic cross-entropy loss (BAS-Loss), defined as: l3=−∑i,cI(bi>t)(si,clog⁡s^i,c)l_3 = -\sum_{i, c} \mathbb{I}(b_i > t) \left( s_{i,c} \log \hat{s}_{i,c} \right) where ii indexes spatial pixels, cc indexes semantic classes, bi∈[0,1]b_i \in [0, 1] is the boundary confidence predicted by the boundary head for pixel ii, t∈[0,1]t \in [0, 1] is a predefined threshold, I(⋅)\mathbb{I}(\cdot) is the indicator function, si,c∈{0,1}s_{i,c} \in \{0, 1\} is the one-hot ground-truth class label, and s^i,c∈(0,1)\hat{s}_{i,c} \in (0, 1) is the predicted class probability.

    The default hyperparameter values are λ0=0.4\lambda_0 = 0.4, λ1=20\lambda_1 = 20, λ2=1.0\lambda_2 = 1.0, λ3=1.0\lambda_3 = 1.0, and boundary threshold t=0.8t = 0.8.

  7. Knowl 7 — Real-Time Semantic Segmentation Benchmark Results on Cityscapes

    data/table

    The table reports accuracy (mean Intersection over Union, mIoU), inference speed (frames per second, FPS), floating-point operations (GFLOPs), and parameter count on the Cityscapes dataset (2048×10242048 \times 1024 resolution). Latencies for models marked with an asterisk were benchmarked on a single NVIDIA RTX 3090 GPU with batch size 1 and batch normalization layers fused into preceding convolutions, without TensorRT acceleration. GFLOPs are computed on full-resolution input.

    Model Val mIoU (%) Test mIoU (%) FPS GPU GFLOPs Params
    MSFNet - 77.1 41.0 RTX 2080Ti 96.8 -
    DF2-Seg1 75.9 74.8 67.2 GTX 1080Ti - -
    DF2-Seg2 76.9 75.3 56.3 GTX 1080Ti - -
    SwiftNetRN-18 75.5 75.4 39.9 GTX 1080Ti 104.0 11.8M
    SwiftNetRN-18 ens - 76.5 18.4 GTX 1080Ti 218.0 24.7M
    CABiNet 76.6 75.9 76.5 RTX 2080Ti 12.0 2.64M
    BiSeNet(Res18) 74.8 74.7 65.5 GTX 1080Ti 55.3 49.0M
    BiSeNetV2-L 75.8 75.3 47.3 GTX 1080Ti 118.5 -
    STDC1-Seg75* 74.5 75.3 74.8 RTX 3090 - -
    STDC2-Seg75* 77.0 76.8 58.2 RTX 3090 - -
    PP-LiteSeg-T2* 76.0 74.9 96.0 RTX 3090 - -
    PP-LiteSeg-B2* 78.2 77.5 68.2 RTX 3090 - -
    HyperSeg-M* 76.2 75.8 59.1 RTX 3090 7.5 10.1M
    HyperSeg-S* 78.2 78.1 45.7 RTX 3090 17.0 10.2M
    SFNet(DF2)* - 77.8 87.6 RTX 3090 - 10.53M
    SFNet(ResNet-18)* - 78.9 30.4 RTX 3090 247.0 12.87M
    SFNet(ResNet-18)†* - 80.4 30.4 RTX 3090 247.0 12.87M
    DDRNet-23-S* 77.8 77.4 108.1 RTX 3090 36.3 5.7M
    DDRNet-23* 79.5 79.4 51.4 RTX 3090 143.1 20.1M
    DDRNet-39* - 80.4 30.8 RTX 3090 281.2 32.3M
    PIDNet-S-Simple 78.8 78.2 100.8 RTX 3090 46.3 7.6M
    PIDNet-S 78.8 78.6 93.2 RTX 3090 47.6 7.6M
    PIDNet-M 80.1 80.1 39.8 RTX 3090 197.4 34.4M
    PIDNet-L 80.9 80.6 31.1 RTX 3090 275.8 36.9M

    Note: † indicates pretraining on extra segmentation datasets.

    PIDNet-L achieves 80.6%80.6\% test mIoU at 31.1 FPS, surpassing DDRNet-39 (80.4%80.4\% at 30.8 FPS) and SFNet(ResNet-18)† (80.4%80.4\% at 30.4 FPS). PIDNet-S achieves 78.6%78.6\% test mIoU at 93.2 FPS, outperforming DDRNet-23-S (77.4%77.4\% at 108.1 FPS) and PP-LiteSeg-B2 (77.5%77.5\% at 68.2 FPS). PIDNet-S-Simple maintains 78.2%78.2\% test mIoU with an inference speed of 100.8 FPS (<10<10 ms latency).

  8. Knowl 8 — Real-Time Semantic Segmentation Benchmark Results on CamVid

    data/table

    The table compares semantic segmentation accuracy (mIoU) and inference speed (FPS) on the CamVid dataset (960×720960 \times 720 resolution, 11 evaluated categories). Models marked with † were pretrained on Cityscapes, and speeds marked with * were evaluated on a single RTX 3090 GPU.

    Model mIoU (%) FPS GPU
    MSFNet 75.4 91.0 GTX 2080Ti
    PP-LiteSeg-T 75.0 154.8 GTX 1080Ti
    TD2-PSP50 76.0 11.0 TITAN X
    BiSeNetV2† 76.7 124.0 GTX 1080Ti
    BiSeNetV2-L† 78.5 33.0 GTX 1080Ti
    HyperSeg-S 78.4 38.0 GTX 1080Ti
    HyperSeg-L 79.1 16.6 GTX 1080Ti
    DDRNet-23-S†* 78.6 182.4 RTX 3090
    DDRNet-23†* 80.6 116.8 RTX 3090
    PIDNet-S† 80.1 153.7 RTX 3090
    PIDNet-S-Wider† 82.0 85.6 RTX 3090

    PIDNet-S achieves 80.1%80.1\% mIoU at 153.7 FPS, exceeding DDRNet-23-S (78.6%78.6\% mIoU at 182.4 FPS) by 1.5%1.5\% mIoU with approximately 1 ms additional latency. PIDNet-S-Wider (doubling PIDNet-S channel dimensions) achieves 82.0%82.0\% mIoU at 85.6 FPS.

  9. Knowl 9 — Ablation Studies of PIDNet Architectural Components and Auxiliary Losses

    empirical result

    Ablation experiments conducted on the Cityscapes validation set validate the design choices of PIDNet:

    1. Integration of ADB and Bag into existing two-branch networks:

      • Equipping BiSeNet (ResNet-18) with ADB and Bag improves validation mIoU from 75.4%75.4\% (63.2 FPS) to 76.7%76.7\% (52.1 FPS).
      • Equipping DDRNet-23 with ADB and Bag improves validation mIoU from 79.5%79.5\% (51.4 FPS) to 80.0%80.0\% (39.2 FPS).
    2. Lateral connections and fusion mechanisms (evaluated on PIDNet-L):

      • Baseline without lateral connections using Add fusion: 78.1%78.1\% mIoU (80.0%80.0\% with ImageNet pretraining).
      • Element-wise Add lateral connections with Add fusion: 79.3%79.3\% mIoU (80.7%80.7\% with ImageNet pretraining).
      • Pag lateral connections with Bag fusion: 80.5%80.5\% mIoU (80.9%80.9\% with ImageNet pretraining).
    3. Context harvesting module (evaluated on PIDNet-S):

      • DAPPM with Light-Bag: 78.8%78.8\% mIoU at 83.7 FPS.
      • PAPPM with Add fusion: 78.4%78.4\% mIoU at 97.8 FPS.
      • PAPPM with Light-Bag: 78.8%78.8\% mIoU at 93.2 FPS (matches DAPPM mIoU while accelerating inference by +9.5+9.5 FPS).
    4. Auxiliary losses and training techniques (evaluated on PIDNet-L):

      • Base cross-entropy loss (l2l_2 only): 78.6%78.6\% mIoU.
      • +l0+ l_0 (auxiliary semantic loss): 78.8%78.8\% mIoU (+0.2%+0.2\%).
      • +l1+ l_1 (boundary binary cross-entropy loss): 79.9%79.9\% mIoU (+1.1%+1.1\%).
      • +l3+ l_3 (boundary-awareness semantic loss BAS-Loss): 80.5%80.5\% mIoU (+0.6%+0.6\%).
      • +OHEM+ \text{OHEM} (Online Hard Example Mining): 80.9%80.9\% mIoU (+0.4%+0.4\%).
  10. Knowl 10 — Sensitivity of PIDNet to Boundary Annotation Quality

    limitation

    PIDNet relies fundamentally on boundary detection in the Derivative branch to steer the Boundary-attention-guided fusion module (Bag) and balance detailed and contextual features. Consequently, the architecture's performance depends heavily on the presence of precise, high-quality pixel-level annotations around object boundaries. On datasets with coarse or imprecise boundary delineations, the boundary loss l1l_1 and boundary-awareness cross-entropy loss l3l_3 offer less accurate guidance, which can constrain the network's boundary delineation accuracy and small object parsing relative to datasets with fine annotations such as Cityscapes.

Coverage note — Omitted the standalone PASCAL Context evaluation table as it is a secondary, non-real-time evaluation included only for generalizability comparison against heavy architectures, while all primary real-time benchmark contributions are captured in the Cityscapes and CamVid knowls.

References

  1. 1.Wangpeng An, Haoqian Wang, Qingyun Sun, Jun Xu, Qionghai Dai, and Lei Zhang. A pid controller approach for stochastic optimization of deep networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 8522–8531, 2018.
  2. 2.Saeid Asgari Taghanaki, Kumar Abhishek, Joseph Paul Cohen, Julien Cohen-Adad, and Ghassan Hamarneh. Deep semantic segmentation of natural and medical images: a review. Artificial Intelligence Review, 54(1):137–178, 2021.
  3. 3.Helon Vicente Hultmann Ayala and Leandro dos Santos Coelho. Tuning of pid controller based on a multiobjective genetic algorithm applied to a robotic manipulator. Expert Systems with Applications, 39(10):8968–8974, 2012.
  4. 4.Vijay Badrinarayanan, Alex Kendall, and Roberto Cipolla. Segnet: A deep convolutional encoder-decoder architecture for image segmentation. IEEE transactions on pattern analysis and machine intelligence, 39(12):2481–2495, 2017.
  5. 5.Gabriel J Brostow, Julien Fauqueur, and Roberto Cipolla. Semantic object classes in video: A high-definition ground truth database. Pattern Recognition Letters, 30(2):88–97, 2009.
  6. 6.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Semantic image segmentation with deep convolutional nets and fully connected crfs. arXiv preprint arXiv:1412.7062, 2014.
  7. 7.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE transactions on pattern analysis and machine intelligence, 40(4):834–848, 2017.
  8. 8.Liang-Chieh Chen, George Papandreou, Florian Schroff, and Hartwig Adam. Rethinking atrous convolution for semantic image segmentation. arXiv preprint arXiv:1706.05587, 2017.
  9. 9.Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. Encoder-decoder with atrous separable convolution for semantic image segmentation. In Proceedings of the European conference on computer vision (ECCV), pages 801–818, 2018.
  10. 10.Wuyang Chen, Xinyu Gong, Xianming Liu, Qian Zhang, Yuan Li, and Zhangyang Wang. Fasterseg: Searching for faster real-time semantic segmentation. arXiv preprint arXiv:1912.10917, 2019.
  11. 11.Franc¸ois Chollet. Xception: Deep learning with depthwise separable convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1251–1258, 2017.
  12. 12.Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3213–3223, 2016.
  13. 13.Ruoxi Deng, Chunhua Shen, Shengjun Liu, Huibing Wang, and Xinru Liu. Learning to predict crisp boundaries. In Proceedings of the European Conference on Computer Vision (ECCV), pages 562–578, 2018.
  14. 14.Henghui Ding, Xudong Jiang, Bing Shuai, Ai Qun Liu, and Gang Wang. Context contrasted feature and gated multiscale aggregation for scene segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2393–2402, 2018.
  15. 15.Mingyuan Fan, Shenqi Lai, Junshi Huang, Xiaoming Wei, Zhenhua Chai, Junfeng Luo, and Xiaolin Wei. Rethinking bisenet for real-time semantic segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9716–9725, 2021.
  16. 16.Di Feng, Christian Haase-Schutz, Lars Rosenbaum, Heinz ¨ Hertlein, Claudius Glaeser, Fabian Timm, Werner Wiesbeck, and Klaus Dietmayer. Deep multi-modal object detection and semantic segmentation for autonomous driving: Datasets, methods, and challenges. IEEE Transactions on Intelligent Transportation Systems, 22(3):1341–1360, 2020.
  17. 17.Jun Fu, Jing Liu, Haijie Tian, Yong Li, Yongjun Bao, Zhiwei Fang, and Hanqing Lu. Dual attention network for scene segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3146–3154, 2019.
  18. 18.Mostafa Gamal, Mennatullah Siam, and Moemen Abdel-Razek. Shuffleseg: Real-time semantic segmentation network. arXiv preprint arXiv:1803.03816, 2018.
  19. 19.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  20. 20.Yuanduo Hong, Huihui Pan, Weichao Sun, and Yisong Jia. Deep dual-resolution networks for real-time and accurate semantic segmentation of road scenes. arXiv preprint arXiv:2101.06085, 2021.
  21. 21.Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861, 2017.
  22. 22.Ping Hu, Fabian Caba, Oliver Wang, Zhe Lin, Stan Sclaroff, and Federico Perazzi. Temporally distributed networks for fast video semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8818–8827, 2020.
  23. 23.Zilong Huang, Xinggang Wang, Lichao Huang, Chang Huang, Yunchao Wei, and Wenyu Liu. Ccnet: Criss-cross attention for semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 603–612, 2019.
  24. 24.A Jayachitra and R Vinodha. Genetic algorithm based pid controller tuning approach for continuous stirred tank reactor. Advances in Artificial Intelligence (16877470), 2014.
  25. 25.A Khodabakhshian and R Hooshmand. A new pid controller design for automatic generation control of hydro power systems. International Journal of Electrical Power & Energy Systems, 32(5):375–382, 2010.
  26. 26.Saumya Kumaar, Ye Lyu, Francesco Nex, and Michael Ying Yang. Cabinet: efficient context aggregation network for low-latency semantic segmentation. In 2021 IEEE International Conference on Robotics and Automation (ICRA), pages 13517–13524. IEEE, 2021.
  27. 27.Hanchao Li, Pengfei Xiong, Haoqiang Fan, and Jian Sun. Dfanet: Deep feature aggregation for real-time semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9522–9531, 2019.
  28. 28.Xiangtai Li, Ansheng You, Zhen Zhu, Houlong Zhao, Maoke Yang, Kuiyuan Yang, Shaohua Tan, and Yunhai Tong. Semantic flow for fast and accurate scene parsing. In European Conference on Computer Vision, pages 775–793. Springer, 2020.
  29. 29.Xin Li, Yiming Zhou, Zheng Pan, and Jiashi Feng. Partial order pruning: for best speed/accuracy trade-off in neural architecture search. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9145–9153, 2019.
  30. 30.Guosheng Lin, Anton Milan, Chunhua Shen, and Ian Reid. Refinenet: Multi-path refinement networks for high-resolution semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1925–1934, 2017.
  31. 31.Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3431–3440, 2015.
  32. 32.Ruijun Ma, Shuyi Li, Bob Zhang, and Zhengming Li. Towards fast and robust real image denoising with attentive neural network and pid controller. IEEE Transactions on Multimedia, 2021.
  33. 33.Roozbeh Mottaghi, Xianjie Chen, Xiaobai Liu, Nam-Gyu Cho, Seong-Whan Lee, Sanja Fidler, Raquel Urtasun, and Alan Yuille. The role of context for object detection and semantic segmentation in the wild. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 891–898, 2014.
  34. 34.Yuval Nirkin, Lior Wolf, and Tal Hassner. Hyperseg: Patch-wise hypernetwork for real-time semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4061–4070, 2021.
  35. 35.Marin Orsic, Ivan Kreso, Petra Bevandic, and Sinisa Segvic. In defense of pre-trained imagenet architectures for real-time semantic segmentation of road-driving images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12607–12616, 2019.
  36. 36.Adam Paszke, Abhishek Chaurasia, Sangpil Kim, and Eugenio Culurciello. Enet: A deep neural network architecture for real-time semantic segmentation. arXiv preprint arXiv:1606.02147, 2016.
  37. 37.Juncai Peng, Yi Liu, Shiyu Tang, Yuying Hao, Lutao Chu, Guowei Chen, Zewu Wu, Zeyu Chen, Zhiliang Yu, Yuning Du, et al. Pp-liteseg: A superior real-time semantic segmentation model. arXiv preprint arXiv:2204.02681, 2022.
  38. 38.Rudra PK Poudel, Ujwal Bonde, Stephan Liwicki, and Christopher Zach. Contextnet: Exploring context and detail for semantic segmentation in real-time. arXiv preprint arXiv:1805.04554, 2018.
  39. 39.Rudra PK Poudel, Stephan Liwicki, and Roberto Cipolla. Fast-scnn: Fast semantic segmentation network. arXiv preprint arXiv:1902.04502, 2019.
  40. 40.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention, pages 234–241. Springer, 2015.
  41. 41.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. International journal of computer vision, 115(3):211–252, 2015.
  42. 42.Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4510–4520, 2018.
  43. 43.Abhinav Shrivastava, Abhinav Gupta, and Ross Girshick. Training region-based object detectors with online hard example mining. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 761–769, 2016.
  44. 44.Alexey A Shvets, Alexander Rakhlin, Alexandr A Kalinin, and Vladimir I Iglovikov. Automatic instrument segmentation in robot-assisted surgery using deep learning. In 2018 17th IEEE International Conference on Machine Learning and Applications (ICMLA), pages 624–628. IEEE, 2018.
  45. 45.Haiyang Si, Zhiqiang Zhang, Feifan Lv, Gang Yu, and Feng Lu. Real-time semantic segmentation via multiply spatial fusion network. arXiv preprint arXiv:1911.07217, 2019.
  46. 46.Towaki Takikawa, David Acuna, Varun Jampani, and Sanja Fidler. Gated-scnn: Gated shape cnns for semantic segmentation. In Proceedings of the IEEE/CVF international conference on computer vision, pages 5229–5238, 2019.
  47. 47.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  48. 48.Jingdong Wang, Ke Sun, Tianheng Cheng, Borui Jiang, Chaorui Deng, Yang Zhao, Dong Liu, Yadong Mu, Mingkui Tan, Xinggang Wang, et al. Deep high-resolution representation learning for visual recognition. IEEE transactions on pattern analysis and machine intelligence, 43(10):3349–3364, 2020.
  49. 49.Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. Non-local neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7794–7803, 2018.
  50. 50.Jiacong Xu and Shankar P Bhattacharyya. A pid controller architecture inspired enhancement to the pso algorithm. In Future of Information and Communication Conference, pages 587–603. Springer, 2022.
  51. 51.Changqian Yu, Changxin Gao, Jingbo Wang, Gang Yu, Chunhua Shen, and Nong Sang. Bisenet v2: Bilateral network with guided aggregation for real-time semantic segmentation. International Journal of Computer Vision, 129(11):3051–3068, 2021.
  52. 52.Changqian Yu, Jingbo Wang, Chao Peng, Changxin Gao, Gang Yu, and Nong Sang. Bisenet: Bilateral segmentation network for real-time semantic segmentation. In Proceedings of the European conference on computer vision (ECCV), pages 325–341, 2018.
  53. 53.Fisher Yu and Vladlen Koltun. Multi-scale context aggregation by dilated convolutions. arXiv preprint arXiv:1511.07122, 2015.
  54. 54.Xiaohui Yuan, Jianfang Shi, and Lichuan Gu. A review of deep learning methods for semantic segmentation of remote sensing imagery. Expert Systems with Applications, 169:114417, 2021.
  55. 55.Yuhui Yuan, Xilin Chen, and Jingdong Wang. Object-contextual representations for semantic segmentation. In European conference on computer vision, pages 173–190. Springer, 2020.
  56. 56.Hang Zhang, Kristin Dana, Jianping Shi, Zhongyue Zhang, Xiaogang Wang, Ambrish Tyagi, and Amit Agrawal. Context encoding for semantic segmentation. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pages 7151–7160, 2018.
  57. 57.Xiangyu Zhang, Xinyu Zhou, Mengxiao Lin, and Jian Sun. Shufflenet: An extremely efficient convolutional neural network for mobile devices. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pages 6848–6856, 2018.
  58. 58.Hengshuang Zhao, Xiaojuan Qi, Xiaoyong Shen, Jianping Shi, and Jiaya Jia. Icnet for real-time semantic segmentation on high-resolution images. In Proceedings of the European conference on computer vision (ECCV), pages 405–420, 2018.
  59. 59.Hengshuang Zhao, Jianping Shi, Xiaojuan Qi, Xiaogang Wang, and Jiaya Jia. Pyramid scene parsing network. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2881–2890, 2017.
  60. 60.Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip HS Torr, et al. Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6881–6890, 2021.

Citation

MLA
Xu, J., et al. “PIDNet: A Real-time Semantic Segmentation Network Inspired by PID Controllers”. arXiv, 2022, http://arxiv.org/abs/2206.02066v3.
APA
Xu, J., Xiong, Z., & Bhattacharyya, S. P. (2022). PIDNet: A Real-time Semantic Segmentation Network Inspired by PID Controllers. arXiv. http://arxiv.org/abs/2206.02066v3
Chicago
Xu, J., Z. Xiong, and S. P. Bhattacharyya. 2022. “PIDNet: A Real-time Semantic Segmentation Network Inspired by PID Controllers”. arXiv. http://arxiv.org/abs/2206.02066v3.
Harvard
Xu, J., Xiong, Z. and Bhattacharyya, S.P. (2022) “PIDNet: A Real-time Semantic Segmentation Network Inspired by PID Controllers”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2206.02066v3.
Vancouver
1. Xu J, Xiong Z, Bhattacharyya SP (2022) PIDNet: A Real-time Semantic Segmentation Network Inspired by PID Controllers. arXiv

BibTeX

@article{xu2022pidnet,
  title = {PIDNet: A Real-time Semantic Segmentation Network Inspired by PID Controllers},
  author = {Xu, Jiacong and Xiong, Zixiang and Bhattacharyya, Shankar P.},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2206.02066v3},
  eprint = {2206.02066}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE