DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting

Yongming RaoWenliang ZhaoGuangyi ChenYansong TangZheng ZhuGuan HuangJie ZhouJiwen Lu

article2022CVPR865 citations

Proposes a model-agnostic framework that adapts pre-trained vision-language models to dense prediction tasks like semantic segmentation and object detection by converting image-text matching into pixel-text score maps guided by context-aware prompting.

Listen

Dense prediction tasks such as semantic segmentation, object detection, and instance segmentation require detailed pixel-level analysis, making them computationally expensive and reliant on costly human annotations. While foundation models pre-trained on paired images and natural language text (such as CLIP) have demonstrated remarkable adaptability across standard image classification benchmarks, their rich linguistic knowledge remains underutilized in dense, fine-grained visual prediction due to the structural gap between image-level language contrast and per-pixel spatial reasoning.

The article introduces and evaluates DenseCLIP, a model-agnostic framework designed to transfer the multi-modal representations of vision-language models into pixel-level prediction pipelines. The core objective is to determine whether incorporating explicit language-guided matching and visual context prompting can boost the accuracy and computational efficiency of standard dense prediction models.

To bridge the gap between image-level concepts and individual pixels, the authors reformulated the contrastive vision-language problem into a pixel-text matching mechanism. This mechanism computes spatial score maps between visual features and language embeddings to explicitly guide downstream task decoders and auxiliary objectives. In addition, the method incorporates context-aware prompting via Transformer modules to refine text embeddings using image context. The framework was evaluated across standard benchmarks, including the ADE20K semantic segmentation dataset and the COCO detection and instance segmentation benchmarks, across standard convolutional and Transformer backbones.

The empirical findings demonstrate notable performance gains across multiple tasks. On the ADE20K segmentation benchmark, a standard ResNet-50 backbone equipped with DenseCLIP achieved a 43.5% mean intersection-over-union score, outperforming conventional ImageNet pre-training by 4.9 percentage points and basic vision-language fine-tuning by 3.9 percentage points. A ResNet-101 model paired with DenseCLIP achieved 46.5% multi-scale segmentation accuracy, surpassing heavier competitive architectures while requiring only about one-third of the computation. The framework also delivered consistent gains on the COCO benchmark, yielding a 2.6 percentage point increase in detection average precision and a 2.9 percentage point increase in instance segmentation mask precision over supervised ImageNet baselines, while also improving standard vision backbones like Swin Transformers by up to 2.6 percentage points.

These results demonstrate that language priors can significantly improve spatial representation learning and holistic object recognition in computer vision systems without adding prohibitive computational overhead. In practice, adopting DenseCLIP enables organizations to deploy lightweight decoders that deliver superior segmentation accuracy at a fraction of the computational expense, offering substantial cost and latency savings for production deployments. DenseCLIP also proves flexible enough to guide arbitrary image backbones, allowing existing visual pipelines to integrate language guidance flexibly.

Engineering teams should consider adopting vision-language fine-tuning with post-model prompting for dense spatial tasks to optimize the trade-off between computational budget and predictive accuracy. Specifically, practitioners should adopt customized fine-tuning hyperparameters—such as AdamW optimization and reduced backbone learning rates—to preserve pre-trained knowledge. Future development should explore integrating dense, localized supervision during initial multi-modal pre-training or refining spatial locality mechanisms to further close the performance gap on object detection tasks.

The primary limitation highlighted in the article is that performance gains in object detection were comparatively modest compared to the substantial improvements in segmentation, as image-level contrastive pre-training naturally lacks fine-grained spatial locality constraints. However, the evaluation across multiple recognized benchmarks and diverse neural network architectures provides high confidence in the framework's effectiveness and practical utility for dense prediction systems.

  • Paper: Learning Transferable Visual Models From Natural Language Supervision, Alec Radford et al.. DenseCLIP adapts CLIP’s shared image–text representations for pixel-level prediction, so CLIP’s contrastive pretraining and zero-shot transfer provide the essential foundation.
  • Paper: Language-driven Semantic Segmentation, Boyi Li et al. (2022). LSeg establishes language embeddings as pixel-level labels for zero-shot segmentation, making its dense text–visual alignment a direct precursor to DenseCLIP’s pixel–text matching.
Cover for DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting

Abstract

Recent progress has shown that large-scale pre-training using contrastive image-text pairs can be a promising alternative for high-quality visual representation learning from natural language supervision. Benefiting from a broader source of supervision, this new paradigm exhibits impressive transferability to downstream classification tasks and datasets. However, the problem of transferring the knowledge learned from image-text pairs to more complex dense prediction tasks has barely been visited. In this work, we present a new framework for dense prediction by implicitly and explicitly leveraging the pre-trained knowledge from CLIP. Specifically, we convert the original image-text matching problem in CLIP to a pixel-text matching problem and use the pixel-text score maps to guide the learning of dense prediction models. By further using the contextual information from the image to prompt the language model, we are able to facilitate our model to better exploit the pre-trained knowledge. Our method is model-agnostic, which can be applied to arbitrary dense prediction systems and various pre-trained visual backbones including both CLIP models and ImageNet pre-trained models. Extensive experiments demonstrate the superior performance of our methods on semantic segmentation, object detection, and instance segmentation tasks. Code is available at https://github.com/raoyongming/DenseCLIP.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Approach
  • 3.1 Preliminaries: Overview of CLIP
  • 3.2 Language-Guided Dense Prediction
  • 3.3 Context-Aware Prompting
  • 3.4 Instantiations
  • 4 Experiments
  • 4.1 Semantic Segmentation
  • 4.2 Object Detection and Instance Segmentation
  • 4.3 DenseCLIP for Any Visual Backbone
  • 4.4 Visualization
  • 5 Conclusion and Discussion
  • References

Knowls

  1. Knowl 1 — DenseCLIP converts CLIP features into pixel–text matches

    model/method

    DenseCLIP repurposes CLIP’s spatial image features and class-name text embeddings to form a language-compatible score for every image location. For a CLIP ResNet, let x4∈RP×Cx_4\in\mathbb{R}^{P\times C} be the flattened feature map from the last stage, where P=H4W4P=H_4W_4 spatial locations and CC channels. CLIP averages x4x_4 spatially to obtain xˉ4\bar{x}_4, then applies its attention-pooling multi-head self-attention layer to the concatenated global and spatial features, producing [zˉ,z]=MHSA⁡([xˉ4,x4])[\bar{z},z]=\operatorname{MHSA}([\bar{x}_4,x_4]). DenseCLIP uses the spatial outputs zz as the pixel features; for a ViT image encoder, the corresponding spatial outputs are obtained by excluding the class token.

    For KK task classes, CLIP’s text encoder maps prompts containing the class names to embeddings t∈RK×Ct\in\mathbb{R}^{K\times C}. Row-wise ℓ2\ell_2 normalization of zz and tt gives z^\hat z and t^\hat t. DenseCLIP computes the pixel–text score map

    s=z^t^⊤∈RP×K.s=\hat z\hat t^{\top}\in\mathbb{R}^{P\times K}.

    Each score is a normalized feature similarity between one image location and one class description. The resulting KK-channel map acts as a low-resolution dense prediction and makes CLIP’s image–text alignment usable at pixel level.

  2. Knowl 2 — Pixel–text scores provide language features to a standard dense decoder

    model/method

    DenseCLIP uses the pixel–text score map s∈RP×Ks\in\mathbb{R}^{P\times K} in two ways: it can be supervised as a low-resolution prediction, and it is concatenated with the last image feature map before dense decoding. If x4∈RP×Cx_4\in\mathbb{R}^{P\times C} is that feature map, the decoder input is x4′=[x4,s]∈RP×(C+K)x'_4=[x_4,s]\in\mathbb{R}^{P\times(C+K)}. A standard segmentation or detection pipeline can then consume the augmented feature map, with input-channel dimensions adjusted as needed (for example, for a feature pyramid network). This design makes the language-derived information an explicit decoder input rather than relying only on CLIP initialization.

  3. Knowl 3 — Context-aware prompting adapts class text features to each image

    model/method

    DenseCLIP proposes two ways to condition language features on visual context. In pre-model prompting, a Transformer decoder receives learnable queries q∈RN×Cq\in\mathbb{R}^{N\times C} and the image features [zˉ,z][\bar z,z], producing visual contexts vpre=TransDecoder⁡(q,[zˉ,z])v_{\mathrm{pre}}=\operatorname{TransDecoder}(q,[\bar z,z]). These contexts replace the learnable textual context tokens in the text encoder input. Because the text-encoder input then depends on the image, this option requires text-encoder computation at inference.

    In post-model prompting, learnable language contexts first produce class embeddings tt (as in CoOp); those embeddings query a Transformer decoder over [zˉ,z][\bar z,z] to produce class-specific visual contexts vpostv_{\mathrm{post}}. The text embeddings are updated by the residual rule t←t+γvpostt\leftarrow t+\gamma v_{\mathrm{post}}, where γ∈RC\gamma\in\mathbb{R}^{C} is learnable and initialized to a small value such as 10−410^{-4} to retain the pretrained language prior. DenseCLIP uses post-model prompting by default: it performed better than pre-model prompting in the reported ablation, and stored text features can be reused at inference without running the text encoder for every image. The reported Transformer decoder has 6 layers and 4 attention heads; the language-only prompt baseline uses 8 context tokens, and the text encoder is frozen during training.

  4. Knowl 4 — Pixel–text maps receive task-specific auxiliary supervision

    equation

    DenseCLIP adds an auxiliary loss directly to the pixel–text scores s∈RP×Ks\in\mathbb{R}^{P\times K}, using temperature τ=0.07\tau=0.07. For semantic segmentation, let y∈{1,…,K}Py\in\{1,\ldots,K\}^{P} be the ground-truth class at each score-map location. The auxiliary loss is

    Lauxseg=CrossEntropy⁡(Softmax⁡(s/τ),y).\mathcal{L}_{\mathrm{aux}}^{\mathrm{seg}}=\operatorname{CrossEntropy}(\operatorname{Softmax}(s/\tau),y).

    For detection and instance segmentation, where pixelwise semantic labels are unavailable, DenseCLIP forms a binary target y~∈{0,1}P×K\tilde y\in\{0,1\}^{P\times K} from the ground-truth boxes and their class labels. The auxiliary loss is

    Lauxdet=BinaryCrossEntropy⁡(Sigmoid⁡(s/τ),y~).\mathcal{L}_{\mathrm{aux}}^{\mathrm{det}}=\operatorname{BinaryCrossEntropy}(\operatorname{Sigmoid}(s/\tau),\tilde y).

    These losses train the pixel–text maps alongside the downstream task objective; the paper states that the segmentation auxiliary objective helps the feature map recover spatial locality.

  5. Knowl 5 — ADE20K segmentation improves across CLIP backbones and pretraining strategies

    empirical result

    DenseCLIP was evaluated with Semantic FPN on ADE20K, which has 150 categories, 20K training images, and 2K validation images. The metric is validation mean intersection over union (mIoU), reported for single-scale (SS) and multi-scale (MS) testing. With ResNet-50, ImageNet-pretrained Semantic FPN scored 38.6/40.6 SS/MS, vanilla CLIP fine-tuning scored 39.6/41.6, and DenseCLIP scored 43.5/44.7. With ResNet-101, the corresponding results were 40.4/42.3, 42.7/44.3, and 45.1/46.5. With ViT-B, ImageNet-pretrained Semantic FPN scored 48.3/50.9, CLIP fine-tuning scored 49.4/50.3, and DenseCLIP scored 50.6/51.3; the ImageNet-21K baseline scored 49.1/50.4. Thus DenseCLIP’s single-scale gains over the ImageNet baselines were 4.9, 4.7, and 2.3 mIoU points for ResNet-50, ResNet-101, and ViT-B, respectively, and its gains over vanilla CLIP fine-tuning were 3.9, 2.4, and 1.2 points.

    The ResNet-101 DenseCLIP result was 46.5 MS mIoU at 346.3 GFLOPs and 67.8M parameters; for comparison, DeepLabV3+ reached 46.1 at 1022.7 GFLOPs and UperNet reached 44.8 at 1031.0 GFLOPs. FLOPs were measured at 1024 × 1024 input resolution. The authors also report that default SGD-based CLIP fine-tuning gave only 21.9 mIoU on ADE20K; their CLIP segmentation experiments instead used AdamW and set the image-encoder learning rate to one tenth of the rate for the other parameters to better preserve pretrained weights.

  6. Knowl 6 — Ablations isolate gains from language and visual-context prompts

    empirical result

    On ADE20K with a ResNet-50 and Semantic FPN, the reported ablation compared mIoU, GFLOPs, and parameter count for different prompting configurations. The ImageNet-pretrained baseline scored 38.6 mIoU with 227 GFLOPs and 31.0M parameters; vanilla CLIP fine-tuning scored 39.6 with 249 GFLOPs and 31.0M parameters. Adding language-domain prompting raised the score to 42.1, with 269 GFLOPs and 46.5M parameters. Adding pre-model vision-to-language prompting to the language prompt gave 42.9, with 368 GFLOPs and 116.9M parameters. Using post-model vision-to-language prompting instead gave the best score, 43.5, with 269 GFLOPs and 50.2M parameters. The comparison supports the authors’ choice of post-model prompting: it had the highest measured mIoU among these prompt variants while requiring substantially less computation and fewer parameters than pre-model prompting.

  7. Knowl 7 — Language-guided fine-tuning benefits multiple forms of pretraining

    empirical result

    The ADE20K ResNet-50 comparison of vanilla fine-tuning and language-guided fine-tuning reports SS/MS mIoU pairs for ImageNet-1K (38.6/40.6 and 40.8/42.5), ImageNet-21K (40.8/42.5 and 42.5/44.7), ImageNet-1K with DenseCLIP guidance (41.0/43.0), MoCoV2 (37.6/39.3), DenseCL (38.7/40.2), and CLIP (39.6/41.6 and 43.5/44.7 for CLIP plus DenseCLIP). The plotted comparison indicates that CLIP initialization alone exceeds ImageNet-1K vanilla fine-tuning, and that language-guided fine-tuning with context-aware prompting raises the CLIP result beyond the ImageNet-21K result. The ImageNet-1K and CLIP guided results also show that the guidance is not restricted to a CLIP image encoder.

  8. Knowl 8 — DenseCLIP improves COCO detection and instance segmentation

    empirical result

    On COCO val2017 (118K training images and 5K validation images), DenseCLIP was tested with RetinaNet and Mask R-CNN using ResNet-50 and ResNet-101. Models were trained for 12 epochs with AdamW and batch size 16; RetinaNet used gradient clipping with maximum ℓ2\ell_2 norm 0.1 to address a large initial loss.

    For RetinaNet, box AP for ImageNet-pretrained, vanilla CLIP, and DenseCLIP models was 36.3, 36.9, and 37.8 with ResNet-50, and 38.5, 40.5, and 41.1 with ResNet-101. DenseCLIP therefore improved over the ImageNet baseline by 1.5 and 2.6 AP points, and over vanilla CLIP fine-tuning by 0.9 and 0.6 points, respectively.

    For Mask R-CNN, the same three model variants achieved box AP of 38.2, 39.3, and 40.2 with ResNet-50, and 40.0, 42.2, and 42.6 with ResNet-101. Their instance-mask AP was 34.7, 36.8, and 37.6 with ResNet-50, and 36.1, 38.9, and 39.6 with ResNet-101. DenseCLIP’s mask AP gains over ImageNet pretraining were 2.9 and 3.5 points, and gains over vanilla CLIP were 0.8 and 0.7 points. The results show improvements in both detection and instance segmentation, with the larger relative gains over ImageNet baselines occurring for mask AP.

  9. Knowl 9 — Frozen CLIP language guidance also improves non-CLIP image backbones

    model/method

    DenseCLIP can replace the CLIP image encoder with another pretrained 2D visual backbone while retaining the pretrained CLIP text encoder as a source of language guidance. The authors freeze the text encoder during training, allowing the image backbone to adapt to the language-derived pixel–text supervision even when its features are not initially aligned with the text features. The text encoder can be removed after training.

    On ADE20K, Semantic FPN with ImageNet-pretrained ResNet-50 improved from 38.6 to 41.0 SS mIoU and from 40.6 to 43.0 MS mIoU when equipped with DenseCLIP. ResNet-101 improved from 40.4/42.3 to 43.0/45.2 SS/MS. With UperNet, Swin-T improved from 44.5/45.8 to 45.4/46.5, and Swin-S from 47.6/49.5 to 48.3/49.7. These experiments support applicability beyond CLIP image encoders, though the paper reports that such models still lag behind DenseCLIP models using CLIP image encoders.

  10. Knowl 10 — Detection gains are limited relative to segmentation, with locality as a proposed explanation

    limitation

    The authors note that DenseCLIP’s improvements on detection were less substantial than its improvements on segmentation. They conjecture that CLIP’s image encoder lacks spatial locality because CLIP pretraining does not impose a dense spatial constraint, while object-centered tasks provide less dense supervision. They suggest that adding dense supervision during pretraining or improving locality recovery after pretraining could address this limitation; these are proposed directions, not demonstrated fixes.

Coverage note — Omitted the qualitative ADE20K visualization and exhaustive COCO per-IoU and object-size breakdowns; they provide supporting examples and finer-grained metrics but do not add a separate central method or conclusion beyond the quantitative results summarized here.

References

  1. 1.Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. Vqa: Visual question answering. In ICCV, pages 2425–2433, 2015.
  2. 2.Tom B. Brown, Benjamin Mann, Nick Ryder, and Melanie Subbiah et al. Language models are few-shot learners. In NeurIPS, 2020.
  3. 3.Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In ICCV, 2021.
  4. 4.Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In CVPR, pages 6299–6308, 2017.
  5. 5.Kai Chen, Jiaqi Wang, Jiangmiao Pang, et al. MMDetection: Open mmlab detection toolbox and benchmark. arXiv preprint arXiv:1906.07155, 2019.
  6. 6.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. TPAMI, 40(4):834–848, 2017.
  7. 7.Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. Encoder-decoder with atrous separable convolution for semantic image segmentation. In ECCV, pages 801–818, 2018.
  8. 8.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey E. Hinton. A simple framework for contrastive learning of visual representations. In ICML, pages 1597–1607, 2020.
  9. 9.Xinlei Chen, Haoqi Fan, Ross B. Girshick, and Kaiming He. Improved baselines with momentum contrastive learning. CoRR, abs/2003.04297, 2020.
  10. 10.MMSegmentation Contributors. Mmsegmentation: Openmmlab semantic segmentation toolbox and benchmark, 2020.
  11. 11.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, pages 248–255, 2009.
  12. 12.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2020.
  13. 13.Peng Gao, Shijie Geng, Renrui Zhang, Teli Ma, Rongyao Fang, Yongfeng Zhang, Hongsheng Li, and Yu Qiao. Clip-adapter: Better vision-language models with feature adapters. arXiv preprint arXiv:2110.04544, 2021.
  14. 14.Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In CVPR, pages 580–587, 2014.
  15. 15.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollar, and Ross Girshick. Masked autoencoders are scalable vision learners. arXiv preprint arXiv:2111.06377, 2021.
  16. 16.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross B. Girshick. Momentum contrast for unsupervised visual representation learning. In CVPR, pages 9726–9735, 2020.
  17. 17.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Girshick. Mask r-cnn. In ICCV, 2017.
  18. 18.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016.
  19. 19.Ronghang Hu, Marcus Rohrbach, and Trevor Darrell. Segmentation from natural language expressions. In ECCV, pages 108–124, 2016.
  20. 20.Zilong Huang, Xinggang Wang, Lichao Huang, Chang Huang, Yunchao Wei, and Wenyu Liu. Ccnet: Criss-cross attention for semantic segmentation. In ICCV, pages 603–612, 2019.
  21. 21.Alexander Kirillov, Ross Girshick, Kaiming He, and Piotr Dollar. Panoptic feature pyramid networks. In CVPR, pages 6399–6408, 2019.
  22. 22.Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby. Big transfer (bit): General visual representation learning. In ECCV, pages 491–507. Springer, 2020.
  23. 23.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In NeurIPS, pages 1097–1105, 2012.
  24. 24.Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L. Berg, Mohit Bansal, and Jingjing Liu. Less is more: Clipbert for video-and-language learning via sparse sampling. In CVPR, pages 7331–7341, 2021.
  25. 25.Tsung-Yi Lin, Piotr Dollar, Ross B. Girshick, Kaiming He, Bharath Hariharan, and Serge J. Belongie. Feature pyramid networks for object detection. In CVPR, pages 936–944, 2017.
  26. 26.Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollar. Focal loss for dense object detection. In ICCV, pages 2980–2988, 2017.
  27. 27.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, pages 740–755, 2014.
  28. 28.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. arXiv preprint arXiv:2107.13586, 2021.
  29. 29.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In ICCV, 2021.
  30. 30.Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In CVPR, pages 3431–3440, 2015.
  31. 31.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  32. 32.Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In NeurIPS, pages 13–23, 2019.
  33. 33.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. arXiv preprint arXiv:2103.00020, 2021.
  34. 34.Yongming Rao, Wenliang Zhao, Zheng Zhu, Jiwen Lu, and Jie Zhou. Global filter networks for image classification. In NeurIPS, 2021.
  35. 35.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. In NeurIPS, pages 91–99, 2015.
  36. 36.Tal Ridnik, Emanuel Ben-Baruch, Asaf Noy, and Lihi Zelnik-Manor. Imagenet-21k pretraining for the masses. arXiv preprint arXiv:2104.10972, 2021.
  37. 37.Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. VL-BERT: pre-training of generic visual-linguistic representations. In ICLR, 2020.
  38. 38.Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta. Revisiting unreasonable effectiveness of data in deep learning era. In ICCV, pages 843–852, 2017.
  39. 39.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. Training data-efficient image transformers & distillation through attention. In ICML, pages 10347–10357, 2021.
  40. 40.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, pages 5998–6008, 2017.
  41. 41.Mengmeng Wang, Jiazheng Xing, and Yong Liu. Actionclip: A new paradigm for video action recognition. arXiv preprint arXiv:2109.08472, 2021.
  42. 42.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. arXiv preprint arXiv:2102.12122, 2021.
  43. 43.Xinlong Wang, Rufeng Zhang, Chunhua Shen, Tao Kong, and Lei Li. Dense contrastive learning for self-supervised visual pre-training. In CVPR, pages 3024–3033, 2021.
  44. 44.Zihao Wang, Xihui Liu, Hongsheng Li, Lu Sheng, Junjie Yan, Xiaogang Wang, and Jing Shao. CAMP: cross-modal adaptive message passing for text-image retrieval. In ICCV, pages 5763–5772, 2019.
  45. 45.Tete Xiao, Yingcheng Liu, Bolei Zhou, Yuning Jiang, and Jian Sun. Unified perceptual parsing for scene understanding. In ECCV, pages 418–434, 2018.
  46. 46.Zhenda Xie, Yutong Lin, Zhuliang Yao, Zheng Zhang, Qi Dai, Yue Cao, and Han Hu. Self-supervised learning with swin transformers. arXiv preprint arXiv:2105.04553, 2021.
  47. 47.Zhenda Xie, Yutong Lin, Zheng Zhang, Yue Cao, Stephen Lin, and Han Hu. Propagate yourself: Exploring pixel-level consistency for unsupervised visual representation learning. In CVPR, pages 16684–16693, 2021.
  48. 48.Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron C. Courville, Ruslan Salakhutdinov, Richard S. Zemel, and Yoshua Bengio. Show, attend and tell: Neural image caption generation with visual attention. In ICML, volume 37, pages 2048–2057, 2015.
  49. 49.Zhao Yang, Yansong Tang, Luca Bertinetto, Hengshuang Zhao, and Philip H.S. Torr. Hierarchical interaction network for video object segmentation from referring expressions. In BMVC, 2021.
  50. 50.Zhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen, Hengshuang Zhao, and Philip H.S. Torr. Lavt: Language-aware vision transformer for referring image segmentation. In CVPR, 2022.
  51. 51.Yuan Yao, Ao Zhang, Zhengyan Zhang, Zhiyuan Liu, Tat-Seng Chua, and Maosong Sun. Cpt: Colorful prompt tuning for pre-trained vision-language models. arXiv preprint arXiv:2109.11797, 2021.
  52. 52.Minghao Yin, Zhuliang Yao, Yue Cao, Xiu Li, Zheng Zhang, Stephen Lin, and Han Hu. Disentangled non-local neural networks. In ECCV, 2020.
  53. 53.Fisher Yu and Vladlen Koltun. Multi-scale context aggregation by dilated convolutions. arXiv preprint arXiv:1511.07122, 2015.
  54. 54.Yuhui Yuan, Xilin Chen, and Jingdong Wang. Object-contextual representations for semantic segmentation. In ECCV, 2020.
  55. 55.Hang Zhang, Kristin Dana, Jianping Shi, Zhongyue Zhang, Xiaogang Wang, Ambrish Tyagi, and Amit Agrawal. Context encoding for semantic segmentation. In CVPR, pages 7151–7160, 2018.
  56. 56.Renrui Zhang, Rongyao Fang, Peng Gao, Wei Zhang, Kunchang Li, Jifeng Dai, Yu Qiao, and Hongsheng Li. Tip-adapter: Training-free clip-adapter for better vision-language modeling. arXiv preprint arXiv:2111.03930, 2021.
  57. 57.Hengshuang Zhao, Jianping Shi, Xiaojuan Qi, Xiaogang Wang, and Jiaya Jia. Pyramid scene parsing network. In CVPR, pages 2881–2890, 2017.
  58. 58.Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip H.S. Torr, and Li Zhang. Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In CVPR, 2021.
  59. 59.Bolei Zhou, Hang Zhao, Xavier Puig, Tete Xiao, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Semantic understanding of scenes through the ade20k dataset. IJCV, 127(3):302–321, 2019.
  60. 60.Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Learning to prompt for vision-language models. arXiv preprint arXiv:2109.01134, 2021.

Citation

MLA
Rao, Y., et al. “DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 18061–70, https://doi.org/10.1109/CVPR52688.2022.01755.
APA
Rao, Y., Zhao, W., Chen, G., Tang, Y., Zhu, Z., Huang, G., Zhou, J., & Lu, J. (2022). DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 18061–18070. https://doi.org/10.1109/CVPR52688.2022.01755
Chicago
Rao, Y., W. Zhao, G. Chen, et al. 2022. “DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 18061–70. https://doi.org/10.1109/CVPR52688.2022.01755.
Harvard
Rao, Y. et al. (2022) “DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting”, 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 18061–18070. Available at: https://doi.org/10.1109/CVPR52688.2022.01755.
Vancouver
1. Rao Y, Zhao W, Chen G, Tang Y, Zhu Z, Huang G, Zhou J, Lu J (2022) DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 18061–18070

BibTeX

@inproceedings{Rao_2022, title={DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting}, url={http://dx.doi.org/10.1109/CVPR52688.2022.01755}, DOI={10.1109/cvpr52688.2022.01755}, booktitle={2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Rao, Yongming and Zhao, Wenliang and Chen, Guangyi and Tang, Yansong and Zhu, Zheng and Huang, Guan and Zhou, Jie and Lu, Jiwen}, year={2022}, month=June, pages={18061–18070} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE