Semantic-Aware Representation Blending for Multi-Label Image Recognition with Partial Labels

Tao PuTianshui ChenHefeng WuLiang Lin

article2022AAAI65 citations

Proposes a semantic-aware representation blending framework that transfers category-specific features across images at both instance and prototype levels to effectively complement missing annotations in multi-label image recognition without requiring pre-trained pseudo-labeling models.

Listen

Multi-label image recognition is essential for complex visual tasks such as scene understanding and attribute analysis, but manually labeling every object in large image collections is prohibitively expensive and time-consuming. While training models on partially labeled images offers a practical solution, conventional approaches struggle because they either discard missing labels or rely on pre-trained models to guess unknown categories. These prior techniques degrade significantly when the proportion of known labels is low, creating a barrier to deploying accurate computer vision systems under tight annotation budgets.

The article demonstrates and evaluates a framework called Semantic-Aware Representation Blending, designed to train multi-label recognition models directly from partially annotated data without requiring pre-trained helper models. The core objective is to show that unknown object labels in one image can be effectively filled in by transferring and blending category-specific visual features from other images or learned category prototypes.

To evaluate this framework, the authors conducted comprehensive experiments across three standard benchmark datasets: Microsoft COCO (80 categories), Visual Genome (a subset of 200 categories), and Pascal VOC 2007 (20 categories). The evaluation simulated real-world annotation constraints by testing models across various proportions of known labels, ranging from 10% to 90%. The framework extracts category-specific features from images using a standard visual backbone and applies two blending techniques: an instance-level blending module that shares representations between pairs of images, and a prototype-level blending module that clusters representative features per category and blends them using contrastive learning to ensure training stability.

The experimental findings show that the proposed framework consistently outperforms existing state-of-the-art approaches across all datasets and annotation levels. Most importantly, performance advantages grow wider as available annotations decrease. When only 10% of labels are known, the framework achieves an accuracy improvement (measured by mean average precision) of 4.6 percentage points on Microsoft COCO, 4.6 percentage points on Visual Genome, and 2.2 percentage points on Pascal VOC compared to leading competitors. On the challenging 200-category Visual Genome dataset, the framework achieves an average accuracy of 45.6%, surpassing the previous best method by 4.1 percentage points. Ablation experiments further demonstrate that while instance blending introduces necessary diversity, prototype blending is critical for stabilizing the learning process.

These findings indicate that computer vision models can achieve high classification performance even with minimal manual labeling. Organizations can significantly reduce data labeling costs and operational timelines without sacrificing predictive quality or taking on the complexity and risk of training secondary pseudo-labeling models. Because the framework blends category-specific features rather than entire images, it effectively handles complex scenes with multiple scattered objects.

Organizations developing vision systems with limited annotation resources should consider adopting category-specific feature blending strategies to optimize model performance. For practitioners looking to implement this approach, the source code has been made publicly available. Further deployment efforts should explore applying this architecture to specialized, domain-specific visual datasets beyond standard academic benchmarks.

Confidence in these findings is high due to the consistent gains observed across multiple recognized benchmarks and varied label proportions. However, decision-makers should note that the evaluation relied on simulating missing labels by randomly dropping ground-truth annotations from fully labeled datasets. Performance in production settings where missing labels follow non-random, biased, or systematic human omissions may introduce additional uncertainties that warrant localized validation.

Cover for Semantic-Aware Representation Blending for Multi-Label Image Recognition with Partial Labels

Abstract

Training the multi-label image recognition models with partial labels, in which merely some labels are known while others are unknown for each image, is a considerably challenging and practical task. To address this task, current algorithms mainly depend on pre-training classification or similarity models to generate pseudo labels for the unknown labels. However, these algorithms depend on sufficient multi-label annotations to train the models, leading to poor performance especially with low known label proportion. In this work, we propose to blend category-specific representation across different images to transfer information of known labels to complement unknown labels, which can get rid of pre-training models and thus does not depend on sufficient annotations. To this end, we design a unified semantic-aware representation blending (SARB) framework that exploits instance-level and prototype-level semantic representation to complement unknown labels by two complementary modules: 1) an instance-level representation blending (ILRB) module blends the representations of the known labels in an image to the representations of the unknown labels in another image to complement these unknown labels. 2) a prototype-level representation blending (PLRB) module learns more stable representation prototypes for each category and blends the representation of unknown labels with the prototypes of corresponding labels to complement these labels. Extensive experiments on the MS-COCO, Visual Genome, Pascal VOC 2007 datasets show that the proposed SARB framework obtains superior performance over current leading competitors on all known label proportion settings, i.e., with the mAP improvement of 4.6%, 4.6%, 2.2% on these three datasets when the known label proportion is 10%. Codes are available at https://github.com/HCPLab-SYSU/HCP-MLR-PL.

Table of Contents

  • Introduction
  • Related Work
  • Semantic-aware Representation Blending Overview
  • Instance-Level Representation Blending
  • Prototype-Level Representation Blending
  • Optimization
  • Experiments
  • Experimental Setting
  • Comparison with the State-of-the-art Algorithms
  • Ablative Studies
  • Analysis of the CSRL Module
  • Contribution of the SARB Module
  • Analysis of the ILRB Module
  • Analysis of the PLRB Module
  • Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Semantic-aware representation blending framework

    model/method

    The proposed SARB framework addresses multi-label image recognition with partial labels, where each category label is encoded as 11 (present), −1-1 (absent), or 00 (unknown). SARB learns category-specific representations and complements unknown labels by blending representations from either known instances or category prototypes. It has two complementary branches: instance-level representation blending (ILRB), which transfers category information from another training image, and prototype-level representation blending (PLRB), which transfers information from a learned category prototype. The original labels and the two complemented label sets jointly supervise training. Unlike pseudo-labeling approaches, SARB does not require a separately pretrained classification or similarity model.

  2. Knowl 2 — Category-specific representation and shared prediction pipeline

    model/method

    For an image InI^n, a backbone extracts a global feature map fnf^n. A category-specific representation learner ϕcsrl\phi_{\mathrm{csrl}} produces one representation fcnf_c^n for each of the CC categories:

    [f1n,f2n,…,fCn]=ϕcsrl(fn).[f_1^n,f_2^n,\ldots,f_C^n]=\phi_{\mathrm{csrl}}(f^n).

    Here, nn indexes the image, c∈{1,…,C}c\in\{1,\ldots,C\} indexes a category, and fcnf_c^n is the representation associated with category cc. The paper implements category-specific representation learning with semantic decoupling in its main experiments, while also evaluating semantic attention as an alternative. A shared prediction function ϕ\phi, consisting of a gated neural network, a linear classifier, and a sigmoid, maps the category-specific representations to probabilities scn∈(0,1)s_c^n\in(0,1):

    [s1n,s2n,…,sCn]=ϕ([f1n,f2n,…,fCn]).[s_1^n,s_2^n,\ldots,s_C^n]=\phi([f_1^n,f_2^n,\ldots,f_C^n]).

    The same classifier is used for the original, ILRB-blended, and PLRB-blended representations.

  3. Knowl 3 — Instance-level representation blending

    model/method

    ILRB transfers information about a category that is known in one image to an image where that category is unknown. Let InI^n and ImI^m be two training images, with category-specific representations fcnf_c^n and fcmf_c^m and partial labels ycn,ycm∈{−1,0,1}y_c^n,y_c^m\in\{-1,0,1\}. For category cc, ILRB blends only when ycn=0y_c^n=0 and ycm=1y_c^m=1:

    f^cn={αcfcn+(1−αc)fcm,ycn=0, ycm=1,fcn,otherwise,\hat f_c^n= \begin{cases} \alpha_c f_c^n+(1-\alpha_c)f_c^m,& y_c^n=0,\ y_c^m=1,\\ f_c^n,&\text{otherwise,} \end{cases} y^cn={1−αc,ycn=0, ycm=1,ycn,otherwise.\hat y_c^n= \begin{cases} 1-\alpha_c,& y_c^n=0,\ y_c^m=1,\\ y_c^n,&\text{otherwise.} \end{cases}

    The coefficient αc\alpha_c is learnable and initialized to 0.50.5. Thus, the source image contributes a known positive category representation and the target image receives a soft positive supervision value. The operation is repeated for all categories and the resulting representations f^cn\hat f_c^n are passed through the shared classifier to obtain complemented predictions s^cn\hat s_c^n.

  4. Knowl 4 — Prototype-level representation blending

    model/method

    PLRB produces more stable category-level augmentation by clustering representations rather than always borrowing an individual instance. For category cc, the method collects the representations fcnf_c^n from all training images whose category-cc label is known positive, and applies K-means clustering to obtain KK prototypes Pc=[pc1,pc2,…,pcK]P_c=[p_c^1,p_c^2,\ldots,p_c^K]. For an image InI^n, one unknown category cc is sampled uniformly from {c:ycn=0}\{c:y_c^n=0\}, and one prototype index kk is sampled uniformly from {1,…,K}\{1,\ldots,K\}. The representation and target for that category are replaced by

    f~cn=βcfcn+(1−βc)pck,y~cn=1−βc.\tilde f_c^n=\beta_c f_c^n+(1-\beta_c)p_c^k, \qquad \tilde y_c^n=1-\beta_c.

    All other category representations and labels remain unchanged. The coefficient βc\beta_c is learnable and initialized to 0.50.5. The blended representations are classified by the same prediction function used for the original and ILRB streams.

  5. Knowl 5 — Contrastive learning for compact category prototypes

    equation

    To make representations belonging to the same category more compact before prototype construction, SARB uses a cosine-based contrastive loss. For training images InI^n and ImI^m and category cc, with category-specific representations fcn,fcmf_c^n,f_c^m and partial labels ycn,ycmy_c^n,y_c^m, the pairwise loss is

    ℓcn,m={1−cosine⁡(fcn,fcm),ycn=1 and ycm=1,1+cosine⁡(fcn,fcm),otherwise.\ell_c^{n,m}= \begin{cases} 1-\operatorname{cosine}(f_c^n,f_c^m),&y_c^n=1\text{ and }y_c^m=1,\\ 1+\operatorname{cosine}(f_c^n,f_c^m),&\text{otherwise.} \end{cases}

    The first case encourages two known positive instances of category cc to have similar representations; the second discourages similarity for all other label-pair conditions. With NN training images and CC categories, the total contrastive loss is

    Lcst=∑n=1N∑m=1N∑c=1Cℓcn,m,L_{\mathrm{cst}}=\sum_{n=1}^{N}\sum_{m=1}^{N}\sum_{c=1}^{C}\ell_c^{n,m},

    where cosine⁡(⋅,⋅)\operatorname{cosine}(\cdot,\cdot) denotes cosine similarity.

  6. Knowl 6 — Joint partial-label and contrastive optimization

    equation

    SARB supervises the original prediction stream and both complemented streams with partial binary cross-entropy. For a label vector y∈{−1,0,1}Cy\in\{-1,0,1\}^C and predicted probabilities s∈(0,1)Cs\in(0,1)^C, the partial loss averages only over known labels:

    ℓ(y,s)=−1∑c=1C∣yc∣∑c=1C[1(yc=1)log⁡sc+1(yc=−1)log⁡(1−sc)],\ell(y,s)=-\frac{1}{\sum_{c=1}^{C}|y_c|}\sum_{c=1}^{C}\left[\mathbf{1}(y_c=1)\log s_c+\mathbf{1}(y_c=-1)\log(1-s_c)\right],

    where 1(⋅)\mathbf{1}(\cdot) is the indicator function. The classification objective over NN training images is

    Lcls=∑n=1N[ℓ(yn,sn)+ℓ(y^n,s^n)+ℓ(y~n,s~n)],L_{\mathrm{cls}}=\sum_{n=1}^{N}\left[\ell(y^n,s^n)+\ell(\hat y^n,\hat s^n)+\ell(\tilde y^n,\tilde s^n)\right],

    where (yn,sn)(y^n,s^n) is the original stream, (y^n,s^n)(\hat y^n,\hat s^n) is the ILRB stream, and (y~n,s~n)(\tilde y^n,\tilde s^n) is the PLRB stream. The final objective combines classification and contrastive learning:

    L=Lcls+λLcst.L=L_{\mathrm{cls}}+\lambda L_{\mathrm{cst}}.

    Because the contrastive loss has a much larger numerical magnitude, the experiments use λ=0.05\lambda=0.05.

  7. Knowl 7 — Training and partial-label benchmark construction

    experimental setup

    The experiments use ResNet-101 as the image backbone, initialized from ImageNet. The parameters of its first 91 layers are fixed, while the remaining layers and newly added layers are trained end-to-end with Adam using batch size 1616, momentum coefficients 0.9990.999 and 0.90.9, weight decay 5×10−45\times10^{-4}, initial learning rate 10−510^{-5}, and 20 training epochs; the learning rate is divided by 10 every 10 epochs. Images are resized to 512×512512\times512, randomly cropped with a crop size selected from {512,448,384,320,256}\{512,448,384,320,256\}, resized to 448×448448\times448, and randomly horizontally flipped. ILRB and PLRB are activated from epoch 5, and category prototypes are recomputed every 5 epochs. Both blending modules are removed at inference, when images are resized to 448×448448\times448.

    The evaluation uses MS-COCO, Visual Genome, and Pascal VOC 2007. MS-COCO has 80 categories with 82,801 training and 40,504 validation images. For Visual Genome, the 200 most frequent categories are selected to form VG-200; 98,249 images are used for training and 10,000 for testing. Pascal VOC 2007 has 20 categories and uses its trainval and test splits. Complete annotations are converted to partial labels by randomly dropping positive and negative labels. Known-label proportions range from 10% to 90%, corresponding to dropping 90% to 10% of the labels. Mean average precision is the primary metric, averaged both per known-label setting and across all settings.

  8. Knowl 8 — State-of-the-art performance across datasets

    data/table

    SARB achieves the best reported average performance across the three partial-label benchmarks. The values below are averages over known-label proportions from 10% through 90%; the comparison baseline is the strongest prior method for each metric in the reported results.

    Could not parse LaTeX table

    On MS-COCO, SARB improves average mAP, OF1, and CF1 over the previous best by 2.3, 2.8, and 2.5 percentage points. On VG-200, the corresponding gains are 4.1, 3.8, and 3.8 points; on Pascal VOC 2007 they are 0.7, 0.5, and 1.1 points. The advantage is larger when fewer labels are known: at a 10% known-label proportion, the reported mAP gains are 4.6 points on MS-COCO, 4.6 points on VG-200, and 2.2 points on Pascal VOC 2007. At a 90% known-label proportion, the MS-COCO and Pascal VOC gains are 1.4 and 0.4 points, respectively.

  9. Knowl 9 — Category-specific blending outperforms generic mixup

    empirical result

    Ablations show that the representation must be category-specific for blending to provide useful information. Average mAP on MS-COCO, VG-200, and Pascal VOC 2007 is respectively 77.677.6, 45.445.4, and 90.690.6 when SARB uses semantic attention, and 77.977.9, 45.645.6, and 90.790.7 when it uses semantic decoupling. Semantic decoupling is therefore selected for the main experiments, with gains over semantic attention of 0.30.3, 0.20.2, and 0.10.1 percentage points.

    Generic position-wise mixup is substantially weaker. Image-pixel mixup obtains average mAPs of 74.374.3, 39.739.7, and 89.789.7, while feature-map mixup obtains 74.174.1, 39.639.6, and 89.689.6 on the same three datasets. The SSGRL baseline without SARB obtains 74.174.1, 39.739.7, and 89.589.5. Relative to SARB with category-specific representations, image mixup loses 3.63.6, 5.95.9, and 1.01.0 points, and feature-map mixup loses 3.83.8, 6.06.0, and 1.11.1 points. These results support blending representations associated with the same semantic category rather than blending whole images or undifferentiated feature maps.

  10. Knowl 10 — Complementary roles of ILRB, PLRB, and learnable mixing

    empirical result

    Each blending module improves the SSGRL baseline, while using both gives the strongest result. Average mAP on MS-COCO, VG-200, and Pascal VOC 2007 is reported as follows:

    Could not parse LaTeX table

    Relative to SSGRL, ILRB alone improves mAP by 3.23.2, 5.25.2, and 0.70.7 points, while PLRB alone improves it by 3.23.2, 5.25.2, and 0.90.9 points. Learning the blend coefficients is beneficial: fixing α\alpha at 0.50.5 reduces the ILRB-only results by 0.40.4, 0.40.4, and 0.40.4 points, and fixing β\beta at 0.50.5 reduces the PLRB-only results by 0.40.4, 0.30.3, and 0.20.2 points. Training-loss curves further show that adding PLRB makes optimization less choppy than using ILRB without PLRB, supporting its intended role of supplying stable prototype-based representations alongside diverse instance-based blends.

Coverage note — No substantial contributed component was omitted; detailed definitions of standard evaluation metrics and supplementary per-setting metric values were not expanded because the paper reports them as supplementary evaluation details rather than new SARB components.

References

  1. 1.Abadal, S.; Jain, A.; Guirado, R.; Lopez-Alonso, J.; and Alarcón, E. 2022. Computing Graph Neural Networks: A Survey from Algorithms to Accelerators. ACM Computing Surveys, 54(9): 191:1–191:38.
  2. 2.Chen, L.; Zhan, W.; Tian, W.; He, Y.; and Zou, Q. 2019a. Deep integration: A multi-label architecture for road scene recognition. IEEE Transactions on Image Processing, 28(10): 4883–4898.
  3. 3.Chen, R.; Chen, T.; Hui, X.; Wu, H.; Li, G.; and Lin, L. 2020a. Knowledge Graph Transfer Network for Few-Shot Recognition. In Proceedings of Thirty-Fourth AAAI Conference on Artificial Intelligence, 10575–10582.
  4. 4.Chen, S.; Xie, G.; Liu, Y.; Peng, Q.; Sun, B.; Li, H.; You, X.; and Shao, L. 2021a. HSVA: Hierarchical Semantic-Visual Adaptation for Zero-Shot Learning. In Proceedings of Thirty-Fifth Conference on Neural Information Processing Systems (NeurIPS).
  5. 5.Chen, T.; Lin, L.; Chen, R.; Wu, Y.; and Luo, X. 2018a. Knowledge-Embedded Representation Learning for Fine-Grained Image Recognition. In IJCAI, 627–634.
  6. 6.Chen, T.; Lin, L.; Hui, X.; Chen, R.; and Wu, H. 2020b. Knowledge-Guided Multi-Label Few-Shot Learning for General Image Recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence.
  7. 7.Chen, T.; Pu, T.; Wu, H.; Xie, Y.; Liu, L.; and Lin, L. 2021b. Cross-domain facial expression recognition: A unified evaluation benchmark and adversarial graph learning. IEEE Transactions on Pattern Analysis and Machine Intelligence.
  8. 8.Chen, T.; Wang, Z.; Li, G.; and Lin, L. 2018b. Recurrent Attentional Reinforcement Learning for Multi-label Image Recognition. In Proceedings of AAAI Conference on Artificial Intelligence, 6730–6737.
  9. 9.Chen, T.; Xu, M.; Hui, X.; Wu, H.; and Lin, L. 2019b. Learning semantic-specific graph representation for multi-label image recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 522–531.
  10. 10.Chen, T.; Yu, W.; Chen, R.; and Lin, L. 2019c. Knowledge-embedded routing network for scene graph generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 6163–6171.
  11. 11.Chen, Z.-M.; Wei, X.-S.; Wang, P.; and Guo, Y. 2019d. Multi-label image recognition with graph convolutional networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 5177–5186.
  12. 12.Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009. Imagenet: A large-scale hierarchical image database. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 248–255.
  13. 13.Durand, T.; Mehrasa, N.; and Mori, G. 2019. Learning a deep convnet for multi-label classification with partial labels. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 647–657.
  14. 14.Everingham, M.; Van Gool, L.; Williams, C. K.; Winn, J.; and Zisserman, A. 2010. The pascal visual object classes (voc) challenge. International Journal of Computer Vision, 88(2): 303–338.
  15. 15.Gao, B.-B.; and Zhou, H.-Y. 2021. Learning to Discover Multi-Class Attentional Regions for Multi-Label Image Recognition. IEEE Transactions on Image Processing.
  16. 16.Guo, H.; Zheng, K.; Fan, X.; Yu, H.; and Wang, S. 2019. Visual attention consistency under image transforms for multi-label image classification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 729–739.
  17. 17.He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 770–778.
  18. 18.Huynh, D.; and Elhamifar, E. 2020. Interactive multi-label CNN learning with partial labels. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 9423–9432.
  19. 19.Joulin, A.; Van Der Maaten, L.; Jabri, A.; and Vasilache, N. 2016. Learning visual features from large weakly supervised data. In ECCV, 67–84.
  20. 20.Kim, J.-H.; Choo, W.; and Song, H. O. 2020. Puzzle mix: Exploiting saliency and local statistics for optimal mixup. In Proceedings of International Conference on Machine Learning (ICML), 5275–5285.
  21. 21.Kingma, D. P.; and Ba, J. 2015. Adam: A Method for Stochastic Optimization. In Proceedings of 3rd International Conference on Learning Representations (ICLR), San Diego, USA, May 7-9, 2015.
  22. 22.Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.-J.; Shamma, D. A.; Bernstein, M.; and Fei-Fei, L. 2016. Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations. arXiv preprint arXiv:1602.07332.
  23. 23.Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollar, P.; and Zitnick, C. L. 2014. Microsoft coco: Common objects in context. In ECCV, 740–755.
  24. 24.Liu, N.; Wu, H.; and Lin, L. 2015. Hierarchical Ensemble of Background Models for PTZ-Based Video Surveillance. IEEE Transactions on Cybernetics, 45(1): 89–102.
  25. 25.Sun, C.; Shrivastava, A.; Singh, S.; and Gupta, A. 2017. Revisiting unreasonable effectiveness of data in deep learning era. In Proceedings of IEEE International Conference on Computer Vision (ICCV), 843–852.
  26. 26.Wang, Z.; Chen, T.; Li, G.; Xu, R.; and Lin, L. 2017. Multi-label Image Recognition by Recurrently Discovering Attentional Regions. In Proceedings of IEEE International Conference on Computer Vision (ICCV), 464–472.
  27. 27.Wei, Y.; Xia, W.; Lin, M.; Huang, J.; Ni, B.; Dong, J.; Zhao, Y.; and Yan, S. 2016. HCP: A flexible CNN framework for multi-label image classification. IEEE Transactions on Pattern Analysis and Machine Intelligence, 38(9): 1901–1907.
  28. 28.Wu, X.; Chen, Q.; Li, W.; Xiao, Y.; and Hu, B. 2020. AdaHGNN: Adaptive Hypergraph Neural Networks for Multi-Label Image Classification. In Proceedings of the 28th ACM International Conference on Multimedia (ACMMM), 284–293.
  29. 29.Ye, J.; He, J.; Peng, X.; Wu, W.; and Qiao, Y. 2020. Attention-driven dynamic graph convolutional network for multi-label image recognition. In ECCV, 649–665.
  30. 30.Yun, S.; Han, D.; Oh, S. J.; Chun, S.; Choe, J.; and Yoo, Y. 2019. Cutmix: Regularization strategy to train strong classifiers with localizable features. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 6023–6032.
  31. 31.Zhang, H.; Cisse, M.; Dauphin, Y. N.; and Lopez-Paz, D. 2017. mixup: Beyond empirical risk minimization. arXiv preprint arXiv:1710.09412.
  32. 32.Zhang, W.; Wang, X. E.; Tang, S.; Shi, H.; Shi, H.; Xiao, J.; Zhuang, Y.; and Wang, W. Y. 2020. Relational graph learning for grounded video description generation. In Proceedings of ACM International Conference on Multimedia (ACMMM), 3807–3828.
  33. 33.Zhu, J.; Liao, S.; Lei, Z.; and Li, S. Z. 2017. Multi-label convolutional neural network based pedestrian attribute classification. Image and Vision Computing, 58: 224–229.

Citation

MLA
Pu, T., et al. “Semantic-Aware Representation Blending for Multi-Label Image Recognition with Partial Labels”. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 2, 2022, pp. 2091–98, https://doi.org/10.1609/AAAI.V36I2.20105.
APA
Pu, T., Chen, T., Wu, H., & Lin, L. (2022). Semantic-Aware Representation Blending for Multi-Label Image Recognition with Partial Labels. Proceedings of the AAAI Conference on Artificial Intelligence, 36(2), 2091–2098. https://doi.org/10.1609/AAAI.V36I2.20105
Chicago
Pu, T., T. Chen, H. Wu, and L. Lin. 2022. “Semantic-Aware Representation Blending for Multi-Label Image Recognition with Partial Labels”. Proceedings of the AAAI Conference on Artificial Intelligence 36 (2): 2091–98. https://doi.org/10.1609/AAAI.V36I2.20105.
Harvard
Pu, T. et al. (2022) “Semantic-Aware Representation Blending for Multi-Label Image Recognition with Partial Labels”, Proceedings of the AAAI Conference on Artificial Intelligence, 36(2), pp. 2091–2098. Available at: https://doi.org/10.1609/AAAI.V36I2.20105.
Vancouver
1. Pu T, Chen T, Wu H, Lin L (2022) Semantic-Aware Representation Blending for Multi-Label Image Recognition with Partial Labels. Proceedings of the AAAI Conference on Artificial Intelligence 36:2091–2098

BibTeX

@article{Pu_2022, title={Semantic-Aware Representation Blending for Multi-Label Image Recognition with Partial Labels}, volume={36}, ISSN={2159-5399}, url={http://dx.doi.org/10.1609/AAAI.V36I2.20105}, DOI={10.1609/aaai.v36i2.20105}, number={2}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, publisher={Association for the Advancement of Artificial Intelligence (AAAI)}, author={Pu, Tao and Chen, Tianshui and Wu, Hefeng and Lin, Liang}, year={2022}, month=June, pages={2091–2098} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF