Vision-Language Pre-Training for Multimodal Aspect-Based Sentiment Analysis

Yan LingJianfei YuRui Xia

article2022ACL136 citations

Proposes a unified generative vision-language pre-training framework tailored for multimodal aspect-based sentiment analysis, introducing task-specific pre-training objectives across text, visual, and multimodal levels to achieve state-of-the-art alignment across three key subtasks.

Listen

Social media content increasingly blends text and imagery, making fine-grained sentiment analysis essential for market research, brand monitoring, and public opinion tracking. Understanding these posts requires identifying specific entities or topics (aspects) and evaluating the sentiment attached to each. However, conventional artificial intelligence systems either evaluate text and images independently or rely on general pre-training models that fail to capture the subtle alignments between visual details and specific opinion words.

The article aims to introduce and evaluate a unified vision-language pre-training framework tailored specifically for fine-grained multimodal sentiment analysis. It demonstrates that integrating task-specific pre-training across visual, textual, and joint representations significantly improves the extraction of aspects and the classification of their corresponding sentiments.

The researchers developed a generative encoder-decoder framework and pre-trained it on the MVSA-Multi dataset of over 17,000 multimodal social media posts. The system incorporates five training tasks: two standard reconstruction tasks for masked text and image regions, and three novel task-specific objectives. These task-specific objectives include extracting textual aspect-opinion pairs using entity recognition and sentiment lexicons, generating visual aspect-opinion pairs from image concepts, and predicting overall post sentiment. The framework was evaluated across three core subtasks: extracting aspect terms, classifying aspect sentiment, and jointly extracting aspects with their sentiments on two benchmark social datasets from 2015 and 2017.

The article establishes several key findings. First, the proposed framework consistently outperformed existing state-of-the-art models on joint aspect and sentiment extraction, achieving performance gains of 2.0 to 2.5 percentage points in F1 score over the strongest baselines. Second, task-specific pre-training delivered substantial gains over standard pre-training methods; for example, adding textual aspect-opinion extraction increased aspect extraction performance by 9.44 percentage points under limited supervision. Third, the benefits of the framework are most pronounced in low-resource environments: when only 200 labeled examples were available, pre-training raised the joint extraction score from under 40% to nearly 52% on the 2015 benchmark, while models trained without pre-training degraded substantially.

These findings demonstrate that domain-specific alignment between images and text is critical for automated sentiment interpretation. For organizations deploying social media analytics, adopting this approach can improve the precision of customer insights while sharply reducing the manual data-labeling costs and timelines required to train high-performing models in new domains.

Organizations seeking to analyze complex multimodal data should transition from separate text-image models to unified, task-specific architectures. Decision-makers should evaluate this framework for low-data deployment scenarios where manual annotation is costly or slow. Before large-scale deployment, practitioners should validate the system on larger and more diverse multi-platform datasets and consider explicitly modeling image-text relationship dynamics. Confidence in the reported performance is high for standard social media post formats, though caution is warranted when applying the system to significantly different domains or image styles.

No sufficiently relevant recommendations were found.

Cover for Vision-Language Pre-Training for Multimodal Aspect-Based Sentiment Analysis

Abstract

As an important task in sentiment analysis, Multimodal Aspect-Based Sentiment Analysis (MABSA) has attracted increasing attention in recent years. However, previous approaches either (i) use separately pre-trained visual and textual models, which ignore the cross-modal alignment or (ii) use vision-language models pre-trained with general pre-training tasks, which are inadequate to identify fine-grained aspects, opinions, and their alignments across modalities. To tackle these limitations, we propose a task-specific Vision-Language Pre-training framework for MABSA (VLP-MABSA), which is a unified multimodal encoder-decoder architecture for all the pre-training and downstream tasks. We further design three types of task-specific pre-training tasks from the language, vision, and multimodal modalities, respectively. Experimental results show that our approach generally outperforms the state-of-the-art approaches on three MABSA subtasks. Further analysis demonstrates the effectiveness of each pre-training task. The source code is publicly released at https://github.com/NUSTM/VLP-MABSA.

Table of Contents

  • Abstract
  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 Feature Extractor
  • 3.2 BART-based Generative Framework
  • 3.3 Pre-training Tasks
  • 3.3.1 Textual Pre-training
  • 3.3.2 Visual Pre-training
  • 3.3.3 Multimodal Pre-training
  • 3.3.4 Full Pre-training Loss
  • 3.4 Downstream Tasks
  • 4 Experiment
  • 4.1 Settings
  • 4.2 Compared Systems
  • 4.3 Main Results
  • 4.4 In-depth Analysis of Pre-training Tasks
  • 4.5 Case Study
  • 5 Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Unified multimodal BART architecture

    model/method

    VLP-MABSA uses one BART-based encoder-decoder for both pre-training and downstream multimodal aspect-based sentiment tasks. A Faster R-CNN detector supplies the 36 highest-confidence image regions; each region has a 2,048-dimensional visual feature projected to the model hidden size. Text tokens are embedded, and the text and image features are concatenated with modality boundary markers before encoding. The encoder is a bidirectional Transformer and the shared decoder is autoregressive; task-specific prefix tokens tell the decoder which output to generate. The implementation initializes both six-layer components from BART-base, whose hidden size is 768. Sharing the architecture lets downstream fine-tuning reuse parameters trained on textual, visual, and multimodal objectives.

  2. Knowl 2 — Textual aspect-opinion extraction

    model/method

    The textual Aspect-Opinion Extraction (AOE) task trains the shared decoder to emit token positions for every aspect term and opinion term in a tweet. Since the pre-training corpus has no such annotations, aspect terms are weakly labeled with a tweet-oriented named-entity recognizer, while opinion terms are found by matching words or phrases against SentiWordNet. For a text of length TT containing MM aspect terms and NN opinion terms, the target sequence lists each aspect's start and end positions, a separator token, each opinion's start and end positions, and an end token. The decoder predicts this sequence autoregressively from the multimodal encoder representation. This makes term extraction a generative index-prediction task rather than token-by-token tagging.

  3. Knowl 3 — Visual aspect-opinion generation

    model/method

    Visual Aspect-Opinion Generation (AOG) treats adjective-noun pairs (ANPs) detected in an image as visual aspect-opinion supervision: the noun represents an aspect and the adjective an opinion. A pre-trained DeepSentiBank detector assigns probabilities to 2,089 predefined ANPs, and the highest-probability pair is used as the target. The shared decoder generates the target ANP tokens autoregressively from the multimodal encoder output. The task therefore trains the model to express fine-grained subjective and objective visual content in language, complementing text-derived aspect and opinion signals.

  4. Knowl 4 — Masked text and image reconstruction objectives

    model/method

    VLP-MABSA includes two general denoising tasks alongside its aspect- and sentiment-specific objectives. In Masked Language Modeling (MLM), 15% of text tokens are randomly masked, and the decoder generates the original text conditioned on the masked text-image input. In Masked Region Modeling (MRM), 15% of image regions are replaced by zero vectors in the encoder input; the decoder receives a marker for each region and an MLP predicts the detector's semantic-class distribution for each masked region. The MRM loss is the sum of KL divergences from the Faster R-CNN class distributions q(vz)q(v_z) to the predicted distributions p(vz)p(v_z) over the ZZ masked regions: LMRM=EX∼D∑z=1ZDKL(q(vz)∥p(vz))L_{MRM}=\mathbb{E}_{X\sim D}\sum_{z=1}^{Z}D_{KL}(q(v_z)\|p(v_z)), where XX is a multimodal training example and DD is the pre-training data distribution. MLM reconstructs textual context, while MRM supervises visual-region semantics.

  5. Knowl 5 — Multimodal sentiment prediction and joint pre-training

    model/method

    Multimodal Sentiment Prediction (MSP) uses the coarse positive, neutral, or negative label for each text-image pair to train the shared decoder as a sentiment classifier. The decoder receives an MSP task prefix and its output is passed through an MLP and softmax; cross-entropy against the pair's annotated sentiment is the MSP loss. The five pre-training losses—MLM, textual AOE, MRM, visual AOG, and MSP—are combined as L=λ1LMLM+λ2LAOE+λ3LMRM+λ4LAOG+λ5LMSPL=\lambda_1L_{MLM}+\lambda_2L_{AOE}+\lambda_3L_{MRM}+\lambda_4L_{AOG}+\lambda_5L_{MSP}. The tasks are optimized alternately. In the reported experiments all five loss weights are set to 1. MSP is the objective whose supervision explicitly spans both modalities and provides sentiment labels.

  6. Knowl 6 — Generative formulation of downstream MABSA tasks

    model/method

    VLP-MABSA represents each downstream task as generation of token positions and, where required, sentiment labels. Joint Multimodal Aspect-Sentiment Analysis (JMASA) emits each aspect's start and end positions followed by its polarity; Multimodal Aspect Term Extraction (MATE) emits aspect start and end positions only; and Multimodal Aspect-oriented Sentiment Classification (MASC) emits aspect positions and polarity. The decoder uses the same index-generation mechanism as AOE, with positive, neutral, negative, and end-of-sequence tokens added to its output vocabulary. During MASC inference, the model is given all gold aspect terms and evaluated on all those aspects, rather than evaluated only on correctly predicted aspects.

  7. Knowl 7 — Pre-training data and weak supervision

    data/table

    Pre-training uses MVSA-Multi, a Twitter image-text corpus with pair-level coarse sentiment annotations. The sentiment labels supervise MSP; AOE labels are derived from tweet NER and SentiWordNet matching, and AOG labels come from the highest-scoring DeepSentiBank ANP. The corpus statistics below show that positive examples dominate the data and that the automatically extracted aspect and opinion terms provide substantially more supervision than the number of image-text pairs.

    Sentiment Image-text pairs Aspects Opinions Words
    Positive 11903 10593 22752 215044
    Neutral 4107 3756 7567 74456
    Negative 1500 1016 2956 25211

    The aspect and opinion counts are outputs of the weak-labeling procedures, not human-provided gold annotations.

  8. Knowl 8 — Benchmark performance across three subtasks

    empirical result

    VLP-MABSA was evaluated on TWITTER-2015 and TWITTER-2017 using BART-base with six-layer encoder and decoder, hidden size 768, learning rate 5×10−55\times10^{-5}, pre-training for 40 epochs, and downstream fine-tuning for 35 epochs. Batch sizes were 64 for pre-training and 16 for fine-tuning. JMASA and MATE are reported by F1; MASC is reported by accuracy and F1. The table compares VLP-MABSA with the strongest directly comparable listed baseline for each task.

    Task and metric Dataset Baseline VLP-MABSA
    JMASA F1 TWITTER-2015 64.1 (JML) 66.6
    JMASA F1 TWITTER-2017 66.0 (JML) 68.0
    MATE F1 TWITTER-2015 82.4 (JML-MATE) 85.7
    MATE F1 TWITTER-2017 91.4 (JML-MATE) 91.7
    MASC accuracy / F1 TWITTER-2015 78.0 / 73.2 (CapTrBERT) 78.6 / 73.8
    MASC accuracy / F1 TWITTER-2017 72.3 / 70.2 (CapTrBERT) 73.8 / 71.8

    VLP-MABSA improves JMASA F1 over JML by 2.5 points on TWITTER-2015 and 2.0 points on TWITTER-2017, and obtains the highest listed MATE F1 on both datasets. For TWITTER-2015 MASC, JML reports accuracy 78.7, but evaluates only on correctly predicted aspects; VLP-MABSA evaluates on all gold aspects, so those accuracy figures are not directly comparable.

  9. Knowl 9 — Effects of pre-training tasks and downstream data size

    empirical result

    An incremental ablation evaluated JMASA F1, MATE F1, and MASC accuracy under full supervision and weak supervision. Weak supervision uses 200 randomly selected downstream training examples. Each row adds the named objective to the objectives in preceding rows; T, V, and MM denote textual, visual, and multimodal pre-training. The results generally rise as objectives are added. In weakly supervised TWITTER-2015 MATE, adding AOE raises F1 from 69.69 to 79.13 (9.44 points); adding MSP after AOG raises MASC accuracy from 59.32 to 62.58. The authors also report that gains from pre-training are larger with fewer downstream examples and become smaller as the sample size increases.

    Setting Added task TWITTER-2015 TWITTER-2017
    JMASA F1 MATE F1 MASC Acc. JMASA F1 MATE F1 MASC Acc.
    Full None 65.31 84.80 76.81 66.10 90.67 72.78
    Full TMLM 65.44 84.91 77.08 66.27 91.00 72.82
    Full TAOE 65.92 85.43 77.48 67.12 91.75 72.89
    Full VMRM 65.94 85.49 77.53 67.15 91.72 73.13
    Full VAOG 66.38 85.73 77.82 67.66 91.77 73.32
    Full MMMSP 66.64 85.66 78.59 68.05 91.73 73.82
    Weak None 39.79 69.33 57.40 49.12 80.48 61.04
    Weak TMLM 40.42 69.69 58.00 49.69 81.26 61.15
    Weak TAOE 46.15 79.13 58.32 52.00 84.60 61.46
    Weak VMRM 46.64 79.49 58.68 52.18 84.47 61.78
    Weak VAOG 47.79 80.94 59.32 53.16 85.04 62.51
    Weak MMMSP 51.71 80.69 62.58 55.38 84.88 64.42

    The ablation supports the paper's analysis that general denoising objectives alone yield smaller changes than task-specific objectives: AOE most strongly improves aspect extraction, AOG contributes further gains, and MSP particularly improves sentiment classification.

  10. Knowl 10 — Stated future directions and scope

    limitation

    The authors characterize VLP-MABSA as an initial unified vision-language pre-training framework for MABSA and identify two directions for further work: applying the approach to a larger pre-training dataset and explicitly modeling relations between the image and text during pre-training. These are stated as future directions; the paper does not quantify their effects or claim that either extension has been tested.

Coverage note — No substantial contributed material was deliberately omitted; the illustrative case-study examples are not reproduced because the benchmark and ablation results capture the paper's general empirical findings more compactly.

References

  1. 1.Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018. Bottom-up and top-down attention for image captioning and visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6077–6086.
  2. 2.Damian Borth, Rongrong Ji, Tao Chen, Thomas Breuel, and Shih-Fu Chang. 2013. Large-scale visual sentiment ontology and detectors using adjective noun pairs. In Proceedings of the 21st ACM international conference on Multimedia, pages 223–232.
  3. 3.Guimin Chen, Yuanhe Tian, and Yan Song. 2020a. Joint aspect extraction and sentiment analysis with directional graph convolutional networks. In Proceedings of the 28th International Conference on Computational Linguistics, COLING 2020, Barcelona, Spain (Online), December 8-13, 2020, pages 272–279. International Committee on Computational Linguistics.
  4. 4.Tao Chen, Damian Borth, Trevor Darrell, and Shih-Fu Chang. 2014. Deepsentibank: Visual sentiment concept classification with deep convolutional neural networks. arXiv preprint arXiv:1410.8586.
  5. 5.Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020b. Uniter: Universal image-text representation learning. In European conference on computer vision, pages 104–120. Springer.
  6. 6.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL, pages 4171–4186.
  7. 7.Andrea Esuli and Fabrizio Sebastiani. 2006. Sentiwordnet: A publicly available lexical resource for opinion mining. In Proceedings of the Fifth International Conference on Language Resources and Evaluation (LREC’06).
  8. 8.Weituo Hao, Chunyuan Li, Xiujun Li, Lawrence Carin, and Jianfeng Gao. 2020. Towards learning a generic agent for vision-and-language navigation via pre-training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13137–13146.
  9. 9.Ruidan He, Wee Sun Lee, Hwee Tou Ng, and Daniel Dahlmeier. 2019. An interactive multi-task learning network for end-to-end aspect-based sentiment analysis. In Proceedings of ACL, pages 504–515.
  10. 10.Minghao Hu, Yuxing Peng, Zhen Huang, Dongsheng Li, and Yiwei Lv. 2019. Open-domain targeted sentiment analysis via span-based extraction and classification. In Proceedings of ACL, pages 537–546.
  11. 11.Xincheng Ju, Dong Zhang, Rong Xiao, Junhui Li, Shoushan Li, Min Zhang, and Guodong Zhou. 2021. Joint multi-modal aspect-sentiment analysis with auxiliary cross-modal relation detection. In Proceedings of EMNLP.
  12. 12.Zaid Khan and Yun Fu. 2021. Exploiting bert for multimodal target sentimentclassification through input space translation. In Proceedings of ACM MM.
  13. 13.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of ACL, pages 7871–7880.
  14. 14.Xin Li, Lidong Bing, Piji Li, and Wai Lam. 2019. A unified model for opinion target extraction and target sentiment prediction. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 6714–6721.
  15. 15.Di Lu, Leonardo Neves, Vitor Carvalho, Ning Zhang, and Heng Ji. 2018. Visual attention model for name tagging in multimodal social media. In Proceedings of ACL, pages 1990–1999.
  16. 16.Jiebo Luo, Damian Borth, and Quanzeng You. 2017. Social multimedia sentiment analysis. In Proceedings of the 25th ACM international conference on Multimedia, pages 1953–1954.
  17. 17.Teng Niu, Shiai Zhu, Lei Pang, and Abdulmotaleb El Saddik. 2016. Sentiment analysis on multi-view social data. In International Conference on Multimedia Modeling, pages 15–27. Springer.
  18. 18.Alan Ritter, Sam Clark, Mausam, and Oren Etzioni. 2011. Named entity recognition in tweets: An experimental study. In EMNLP.
  19. 19.Lin Sun, Jiquan Wang, Yindu Su, Fangsheng Weng, Yuxuan Sun, Zengwei Zheng, and Yuanyi Chen. 2020. Riva: A pre-trained tweet multimodal model based on text-image relation for multimodal ner. In Proceedings of the 28th International Conference on Computational Linguistics, pages 1852–1862.
  20. 20.Lin Sun, Jiquan Wang, Kai Zhang, Yindu Su, and Fangsheng Weng. 2021. Rpbert: A text-image relation propagation-based bert model for multimodal ner. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 13860–13868.
  21. 21.Hanqian Wu, Siliang Cheng, Jingjing Wang, Shoushan Li, and Lian Chi. 2020a. Multimodal aspect extraction with region-aware alignment network. In Proceedings of NLPCC.
  22. 22.Zhiwei Wu, Changmeng Zheng, Cai Yi, Leung Ho-fung Chen Junying, and Qing Li. 2020b. Multimodal representation with embedded visual guiding objects for named entity recognition in social media posts. In Proceedings of ACM MM.
  23. 23.Yiran Xing, Zai Shi, Zhao Meng, Gerhard Lakemeyer, Yunpu Ma, and Roger Wattenhofer. 2021. Kmbart: Knowledge enhanced multimodal bart for visual commonsense generation. In Proceedings of ACL.
  24. 24.Nan Xu, Wenji Mao, and Guandan Chen. 2018. A co-memory network for multimodal sentiment analysis. In The 41st international ACM SIGIR conference on research & development in information retrieval, pages 929–932.
  25. 25.Nan Xu, Wenji Mao, and Guandan Chen. 2019. Multi-interactive memory network for aspect based multimodal sentiment analysis. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 371–378.
  26. 26.Ashima Yadav and Dinesh Kumar Vishwakarma. 2020. Sentiment analysis using deep learning architectures: a review. Artificial Intelligence Review, 53(6):4335–4385.
  27. 27.Hang Yan, Junqi Dai, Xipeng Qiu, and Zheng Zhang. 2021. A unified generative framework for aspect-based sentiment analysis. In Proceedings of ACL-IJCNLP, pages 2416–2429.
  28. 28.Li Yang, Jianfei Yu, Chengzhi Zhang, and Jin-Cheon Na. 2021a. Fine-grained sentiment analysis of political tweets with entity-aware multimodal network. In International Conference on Information, pages 411–420.
  29. 29.Xiaocui Yang, Shi Feng, Yifei Zhang, and Daling Wang. 2021b. Multimodal sentiment detection based on multi-channel graph neural networks. In Proceedings of ACL-IJCNLP, pages 328–339.
  30. 30.Quanzeng You, Jiebo Luo, Hailin Jin, and Jianchao Yang. 2015. Joint visual-textual sentiment analysis with deep neural networks. In Proceedings of the 23rd ACM international conference on Multimedia, pages 1071–1074.
  31. 31.Quanzeng You, Jiebo Luo, Hailin Jin, and Jianchao Yang. 2016. Cross-modality consistent regression for joint visual-textual sentiment analysis of social multimedia. In Proceedings of the Ninth ACM international conference on Web search and data mining, pages 13–22.
  32. 32.Fei Yu, Jiji Tang, Weichong Yin, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang. 2021. Ernie-vil: Knowledge enhanced vision-language representations through scene graphs. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 3208–3216.
  33. 33.Jianfei Yu and Jing Jiang. 2019. Adapting BERT for target-oriented multimodal sentiment classification. In Proceedings of IJCAI, pages 5408–5414.
  34. 34.Jianfei Yu, Jing Jiang, and Rui Xia. 2020a. Entity-sensitive attention and fusion network for entity-level multimodal sentiment classification. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 28:429–439.
  35. 35.Jianfei Yu, Jing Jiang, Li Yang, and Rui Xia. 2020b. Improving multimodal named entity recognition via entity span detection with unified multimodal transformer. In Proceedings of ACL.
  36. 36.Dong Zhang, Suzhong Wei, Shoushan Li, Hanqian Wu, Qiaoming Zhu, and Guodong Zhou. 2021a. Multimodal graph fusion for named entity recognition with targeted visual guidance. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 14347–14355.
  37. 37.Meishan Zhang, Yue Zhang, and Duy-Tin Vo. 2015. Neural networks for open domain targeted sentiment. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 612–621.
  38. 38.Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao. 2021b. Vinvl: Revisiting visual representations in vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5579–5588.
  39. 39.Qi Zhang, Jinlan Fu, Xiaoyu Liu, and Xuanjing Huang. 2018. Adaptive co-attention network for named entity recognition in tweets. In Thirty-Second AAAI Conference on Artificial Intelligence.

Citation

MLA
Ling, Y., et al. “Vision-Language Pre-Training for Multimodal Aspect-Based Sentiment Analysis”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 2149–59, https://doi.org/10.18653/v1/2022.acl-long.152.
APA
Ling, Y., Yu, J., & Xia, R. (2022). Vision-Language Pre-Training for Multimodal Aspect-Based Sentiment Analysis. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2149–2159. https://doi.org/10.18653/v1/2022.acl-long.152
Chicago
Ling, Y., J. Yu, and R. Xia. 2022. “Vision-Language Pre-Training for Multimodal Aspect-Based Sentiment Analysis”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2149–59. https://doi.org/10.18653/v1/2022.acl-long.152.
Harvard
Ling, Y., Yu, J. and Xia, R. (2022) “Vision-Language Pre-Training for Multimodal Aspect-Based Sentiment Analysis”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 2149–2159. Available at: https://doi.org/10.18653/v1/2022.acl-long.152.
Vancouver
1. Ling Y, Yu J, Xia R (2022) Vision-Language Pre-Training for Multimodal Aspect-Based Sentiment Analysis. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 2149–2159

BibTeX

@inproceedings{ling-etal-2022-vision,
    title = "Vision-Language Pre-Training for Multimodal Aspect-Based Sentiment Analysis",
    author = "Ling, Yan  and
      Yu, Jianfei  and
      Xia, Rui",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.152/",
    doi = "10.18653/v1/2022.acl-long.152",
    pages = "2149--2159"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/