UniMS: A Unified Framework for Multimodal Summarization with Knowledge Distillation

Zhengkun ZhangXiaojun MengYasheng WangXin JiangQun LiuZhenglu Yang

article2022AAAI61 citations

Proposes a unified multimodal summarization framework extending BART that jointly handles extractive, abstractive, and image selection tasks while using CLIP-based knowledge distillation to eliminate the need for image captions.

Listen

As digital media expands rapidly, users increasingly demand concise pictorial summaries that combine condensed text with the most relevant images. Traditional automated systems typically perform only one type of text summary—either extracting key sentences directly or generating new abstractive phrasing—and rely heavily on high-quality image captions to select pictures. However, real-world multimedia content frequently lacks accurate captions, limiting the practical utility of existing tools.

The article introduces and evaluates UniMS, a unified artificial intelligence framework built on the BART language model. Its primary objective is to simultaneously generate extractive summaries, produce abstractive summaries, and select the most relevant images without needing image captions.

To evaluate this framework, the authors conducted comprehensive experiments on the benchmark MSMO dataset, which contains over 300,000 news articles paired with multiple images. The UniMS approach incorporates linear image patch projections to process visual data efficiently and transfers visual-text matching intelligence from a pretrained vision-language model (CLIP) using knowledge distillation. Furthermore, it incorporates an extractive text reference into the encoder and uses a visually guided decoder to integrate textual and visual signals seamlessly during abstractive text generation.

The experimental and human evaluations established several key findings. First, UniMS established new state-of-the-art performance across all multimodal summarization subtasks, achieving superior text overlap scores and outperforming prior methods in image selection precision (reaching 69.38% precision compared to earlier benchmarks around 59–65%). Second, knowledge distillation proved just as effective as caption-dependent methods, removing the need for manual captions while maintaining high text-image relevance. Third, using lightweight linear patch projections achieved top performance while adding only about 10% parameter overhead, compared to 28–74% extra overhead incurred by complex visual processing backbones. Finally, human evaluation confirmed that UniMS produced significantly more consistent, relevant, and factual summaries than baseline systems.

These results demonstrate that organizations can deploy unified multimodal summarization to process multimedia documents faster and at lower computational costs. By eliminating reliance on image captions and heavy vision backbones, the approach reduces operational overhead and broadens applicability across uncurated multimedia feeds without compromising factual accuracy or relevance.

Decision-makers and technical teams should consider adopting unified encoder-decoder frameworks for multimedia intelligence pipelines, utilizing knowledge distillation to bypass expensive metadata curation like manual captioning. Future technical work should explore pretraining vision-language base models specifically tailored for multimodal summarization tasks to boost performance even further.

Confidence in these findings is high given the robust evaluations on a large standard dataset and rigorous human assessments. A noted limitation is that the model was evaluated primarily on news-style articles with limited visual complexity per story; performance on domain-specific corpora, video inputs, or highly technical documents may require additional validation and fine-tuning.

arXiv: 2109.05812
Cover for UniMS: A Unified Framework for Multimodal Summarization with Knowledge Distillation

Abstract

With the rapid increase of multimedia data, a large body of literature has emerged to work on multimodal summarization, the majority of which target at refining salient information from textual and visual modalities to output a pictorial summary with the most relevant images. Existing methods mostly focus on either extractive or abstractive summarization and rely on qualified image captions to build image references. We are the first to propose a Unified framework for Multimodal Summarization grounding on BART, UniMS, that integrates extractive and abstractive objectives, as well as selecting the image output. Specially, we adopt knowledge distillation from a vision-language pretrained model to improve image selection, which avoids any requirement on the existence and quality of image captions. Besides, we introduce a visual guided decoder to better integrate textual and visual modalities in guiding abstractive text generation. Results show that our best model achieves a new state-of-the-art result on a large-scale benchmark dataset. The newly involved extractive objective as well as the knowledge distillation technique are proven to bring a noticeable improvement to the multimodal summarization task.

Table of Contents

  • Introduction
  • Related Work
  • Methodology
  • Model Overview
  • Multimodal Reference Enhanced Encoder
  • Visual Guided Decoder
  • Experiments
  • Datasets
  • Implementation Details
  • Baselines
  • Automatic Evaluation
  • Model Analysis
  • Conclusion
  • References

Knowls

  1. Knowl 1 — Multimodal Encoder Architecture with Linear Patch Projection in UniMS

    model/method

    The multimodal encoder of UniMS extends the Bidirectional and Auto-Regressive Transformers (BART) encoder to consume both textual and visual modalities without relying on a heavyweight visual backbone.

    Given a multimodal input document D={T,V}D = \{T, V\}, where T={t1,t2,…,tM}T = \{t_1, t_2, \dots, t_M\} is a sequence of MM text tokens and V={v1,v2,…,vN}V = \{v_1, v_2, \dots, v_N\} is a collection of NN images, each image viv_i is sliced into non-overlapping patches and flattened as a sequence {vik}\{v_{ik}\}. The visual tokens are mapped to visual embeddings evie_{v_i} via a learnable linear projection WvW_v, adding a learned visual position embedding vposv_{\text{pos}} and a special leading token v[CLS]v_{[\text{CLS}]}:

    evi=[v[CLS];vi1;… ;vik]Wv+vpose_{v_i} = [v_{[\text{CLS}]}; v_{i1}; \dots; v_{ik}] W_v + v_{\text{pos}}

    Text tokens are augmented with boundary markers t[CLS]t_{[\text{CLS}]} and t[SEP]t_{[\text{SEP}]} and mapped via text projection matrix WtW_t:

    et=[t[CLS];t1;… ;tM;t[SEP]]Wte_t = [t_{[\text{CLS}]}; t_1; \dots; t_M; t_{[\text{SEP}]}] W_t

    The full input embedding sequence is the concatenation of textual and visual sequences with multimodal position embeddings epose_{\text{pos}}:

    e=[et;ev1;… ;evN]+epose = [e_t; e_{v_1}; \dots; e_{v_N}] + e_{\text{pos}}

    The encoder function fencf_{\text{enc}} maps ee to joint contextual representations h=[ht;hv]=fenc(e)h = [h_t; h_v] = f_{\text{enc}}(e), where hth_t and hvh_v represent text and visual contextual hidden states, respectively. For each image viv_i, the hidden state of its leading token v[CLS]v_{[\text{CLS}]} serves as the complete image representation, denoted h⃗v=hv[CLS]\vec{h}_v = h_{v_{[\text{CLS}]}}.

  2. Knowl 2 — Knowledge Distillation for Caption-Free Image Selection in Multimodal Summarization

    model/method

    To train an image selection module without requiring reference image annotations or high-quality image captions, UniMS uses a knowledge distillation framework where a pretrained vision-language model (CLIP) acts as the teacher network and the UniMS encoder acts as the student network.

    For a candidate image v∈Vv \in V and the reference summary text YtY_t, the teacher network CLIP computes cosine similarity between summary text embedding T(Yt)\mathcal{T}(Y_t) and image embedding V(v)\mathcal{V}(v):

    f(v,Yt)=sim(V(v),T(Yt))f(v, Y_t) = \text{sim}(\mathcal{V}(v), \mathcal{T}(Y_t))

    The student network predicts an image relevance score using a linear layer applied to the encoder image representation h⃗v=hv[CLS]\vec{h}_v = h_{v_{[\text{CLS}]}}:

    gv(h⃗v)=Wvh⃗v+bvg_v(\vec{h}_v) = W_v \vec{h}_v + b_v

    where WvW_v and bvb_v are learnable weight and bias parameters.

    Both scores are converted into probability distributions over candidate images VV using a softmax function with temperature τ=10\tau = 10:

    pv(h⃗v,τ)=exp⁡(gv(h⃗v)/τ)∑v′∈Vexp⁡(gv(h⃗v′)/τ)p_v(\vec{h}_v, \tau) = \frac{\exp(g_v(\vec{h}_v)/\tau)}{\sum_{v' \in V} \exp(g_v(\vec{h}_{v'})/\tau)}

    qv(v,Yt,τ)=exp⁡(f(v,Yt)/τ)∑v′∈Vexp⁡(f(v′,Yt)/τ)q_v(v, Y_t, \tau) = \frac{\exp(f(v, Y_t)/\tau)}{\sum_{v' \in V} \exp(f(v', Y_t)/\tau)}

    The student is trained to minimize the Kullback-Leibler (KL) divergence between student distribution pvp_v and teacher distribution qvq_v:

    LKD=LKL(p∥q)=−∑v∈Vpv(h⃗v,τ)ln⁡qv(v,Yt,τ)pv(h⃗v,τ)\mathcal{L}^{\text{KD}} = \mathcal{L}_{\text{KL}}(p \parallel q) = -\sum_{v \in V} p_v(\vec{h}_v, \tau) \ln \frac{q_v(v, Y_t, \tau)}{p_v(\vec{h}_v, \tau)}

  3. Knowl 3 — Extractive Text Supervision via Greedy Oracle in the Multimodal Encoder

    model/method

    To alleviate modality bias and enhance encoder representations for joint multi-task summarization, UniMS incorporates sentence-level extractive text supervision directly into the multimodal encoder.

    In the input text, individual sentences are demarcated with special boundary tokens [CLS][\text{CLS}] and [SEP][\text{SEP}]. The contextual hidden state of the [CLS][\text{CLS}] token preceding sentence tt, denoted h⃗t=ht[CLS]\vec{h}_t = h_{t_{[\text{CLS}]}}, serves as the representation for sentence tt. A linear classification layer predicts an extractive salience score:

    gt(h⃗t)=Wth⃗t+btg_t(\vec{h}_t) = W_t \vec{h}_t + b_t

    where WtW_t and btb_t are learnable parameters. The predicted probability distribution across candidate sentences is:

    pt(h⃗t)=exp⁡(gt(h⃗t))∑t′exp⁡(gt(h⃗t′))p_t(\vec{h}_t) = \frac{\exp(g_t(\vec{h}_t))}{\sum_{t'} \exp(g_t(\vec{h}_{t'}))}

    Oracle extractive target sentences are generated per document using a greedy sentence-selection algorithm that maximizes the ROUGE-L score against the ground-truth abstractive reference summary. The extractive training loss is computed as:

    LExt=−∑t∈Oraclelog⁡pt(h⃗t)\mathcal{L}^{\text{Ext}} = -\sum_{t \in \text{Oracle}} \log p_t(\vec{h}_t)

  4. Knowl 4 — Visual-Guided Cross-Attention Decoder for Multimodal Abstractive Summarization

    model/method

    The UniMS decoder integrates visual and textual encoder outputs through sequential cross-attention layers instead of standard single-layer cross-attention.

    Within each Transformer decoder block, after token self-attention, the intermediate representation yy first attends to the encoded visual representations hvh_v to generate a visual-guided cross-modal feature, and then attends to the textual representations hth_t to produce the final cross-modal hidden state:

    y←LN(y+CROSSATTN(y;hv))y \leftarrow \text{LN}(y + \text{CROSSATTN}(y; h_v))

    y←LN(y+CROSSATTN(y;ht))y \leftarrow \text{LN}(y + \text{CROSSATTN}(y; h_t))

    where LN(⋅)\text{LN}(\cdot) denotes Layer Normalization and CROSSATTN(query;key/value)\text{CROSSATTN}(\text{query}; \text{key/value}) represents multi-head cross-attention.

    Given previous tokens y<jy_{<j} and multimodal encoder output h=[ht;hv]h = [h_t; h_v], the decoder outputs next-token conditional probabilities py(yj∣y<j,h)p_y(y_j \mid y_{<j}, h). Multimodal abstractive generation is trained by minimizing the autoregressive negative log-likelihood:

    LAbs=−∑j=1∣y∣log⁡py(yj∣y<j,h)\mathcal{L}^{\text{Abs}} = -\sum_{j=1}^{|y|} \log p_y(y_j \mid y_{<j}, h)

  5. Knowl 5 — Unified Multi-Task Training Objective of UniMS

    equation

    The joint optimization objective L\mathcal{L} of the UniMS multimodal summarization framework is the unweighted sum of three loss components across subtasks:

    L=LKD+LExt+LAbs\mathcal{L} = \mathcal{L}^{\text{KD}} + \mathcal{L}^{\text{Ext}} + \mathcal{L}^{\text{Abs}}

    where:

    • LKD=−∑v∈Vpv(h⃗v,τ)ln⁡qv(v,Yt,τ)pv(h⃗v,τ)\mathcal{L}^{\text{KD}} = -\sum_{v \in V} p_v(\vec{h}_v, \tau) \ln \frac{q_v(v, Y_t, \tau)}{p_v(\vec{h}_v, \tau)} is the knowledge distillation loss aligning image selection scores pvp_v from the encoder with CLIP teacher similarity scores qvq_v using temperature τ=10\tau = 10;
    • LExt=−∑t∈Oraclelog⁡pt(h⃗t)\mathcal{L}^{\text{Ext}} = -\sum_{t \in \text{Oracle}} \log p_t(\vec{h}_t) is the extractive sentence selection cross-entropy loss over oracle-labeled document sentences;
    • LAbs=−∑j=1∣y∣log⁡py(yj∣y<j,h)\mathcal{L}^{\text{Abs}} = -\sum_{j=1}^{|y|} \log p_y(y_j \mid y_{<j}, h) is the negative log-likelihood loss for autoregressive abstractive text generation given encoder representations hh.
  6. Knowl 6 — Multimodal Abstractive Summarization and Image Selection Performance on MSMO Dataset

    data/table

    The table below reports test set performance on the MSMO dataset (10,261 test pairs) for abstractive text summarization and image selection. Evaluation metrics include ROUGE-1 (R-1), ROUGE-2 (R-2), ROUGE-L (R-L), Image Precision (IP, in %), and Image-Text relevance (MsimM_{\text{sim}}, in %).

    Model R-1 R-2 R-L IP Msim_{\text{sim}}
    Text Abstractive
    BertAbs (2019) 39.02 18.17 33.20 - -
    BertExtAbs (2019) 39.88 18.77 38.36 - -
    BART (2020) 41.83 19.83 39.74 - -
    Multimodal Abstractive
    ATG (2018) 40.63 18.12 37.53 59.28 25.82
    ATL (2018) 40.86 18.27 37.75 62.44 13.26
    HAN (2018) 40.82 18.30 37.70 61.83 12.22
    MOFencRR_{\text{enc}}^{\text{RR}} (2020) 41.05 18.29 37.74 62.63 26.23
    MOFdecRR_{\text{dec}}^{\text{RR}} (2020) 41.20 18.33 37.80 65.45 26.38
    UniMS 42.94 20.50 40.96 69.38 29.72
    w/o Visual Guide 42.71 20.26 40.76 69.14 29.64
    w/o LExt\mathcal{L}^{\text{Ext}} 42.63 20.21 40.61 69.25 29.61
    w/o Both 42.55 19.88 40.14 69.22 29.57
    UniMS-VL 40.81 18.83 38.83 68.26 29.45

    UniMS achieves superior performance across all text and image metrics. Ablations show that removing visual-guided cross-attention (w/o Visual Guide), extractive supervision (w/o LExt\mathcal{L}^{\text{Ext}}), or both decreases both ROUGE scores and image selection precision (IP).

  7. Knowl 7 — Extractive Summarization Performance on MSMO Dataset

    data/table

    The table below compares extractive summarization performance on the MSMO test set. UniMS performs extractive summarization by ranking document sentences according to their encoder-predicted extractive scores gt(h⃗t)g_t(\vec{h}_t) and selecting the top-3 sentences.

    Model R-1 R-2 R-L
    ORACLE 50.15 28.56 47.91
    LEAD-3 39.94 18.56 38.38
    GR (2018) 37.13 15.03 30.21
    BertExt (2019) 39.02 18.17 33.20
    LAMS-ATL (2021) 42.48 19.75 38.78
    LAMS-MFB (2021) 43.07 20.28 39.34
    UniMS 42.58 20.29 40.91
    UniMS-VL 41.29 19.01 39.47

    UniMS achieves the highest ROUGE-L score (40.91) among all compared extractive baselines, surpassing specialized multimodal extractive models such as LAMS-MFB (39.34 R-L) without requiring explicit image location inputs.

  8. Knowl 8 — Impact of Visual Backbones on Performance and Parameter Overhead

    data/table

    The table below compares the performance and parameter growth of UniMS when employing direct linear patch projection versus various pretrained visual backbones to extract 49 image feature representations on the MSMO dataset.

    Model R-L IP Msim_{\text{sim}} Para.↑\uparrow%
    LinearProjection 40.96 69.38 29.72 10.07
    ResNet50 40.79 68.92 29.66 30.22
    CLIP-ResNet50 40.83 69.14 23.73 28.78
    CLIP-ViT-B-32 40.86 69.16 29.62 74.10
    Fast-RCNN 40.76 68.33 29.62 41.73
    UniMS-VL 38.83 68.26 29.45 41.73

    Direct linear projection on image patches outperforms grid-based features (ResNet50, CLIP-ResNet50, CLIP-ViT-B-32) and region-based features (Faster R-CNN) in ROUGE-L (40.96 vs. 40.76--40.86), IP (69.38 vs. 68.33--69.16), and MsimM_{\text{sim}} (29.72 vs. 23.73--29.66), while adding only 10.07% extra parameters over the text-only BART model compared to 28.78%--74.10% for pretrained visual backbones.

  9. Knowl 9 — Knowledge Distillation versus Caption ROUGE-Ranking for Image Selection Supervision

    data/table

    The table below compares the performance on the MSMO dataset when UniMS is trained with image selection targets derived from CLIP knowledge distillation (KD) versus ROUGE-ranking (RR) based on text-to-caption overlap.

    Image Reference Strategy R-L IP Msim_{\text{sim}}
    ROUGE-ranking 40.90 69.25 29.56
    UniMS (CLIP KD) 40.96 69.38 29.72

    Knowledge distillation from CLIP achieves slightly higher performance than ROUGE-ranking across text ROUGE-L (40.96 vs. 40.90), image selection precision IP (69.38 vs. 69.25), and cross-modal similarity MsimM_{\text{sim}} (29.72 vs. 29.56), while removing the dependency on the availability and quality of image captions.

  10. Knowl 10 — Human Evaluation of Multimodal Summarization Quality

    data/table

    The table below presents human evaluation scores on 200 randomly sampled MSMO test instances assessed by three crowd workers on a 1--3 point Likert scale. Dimensions include Abstractive Consistency (Consist., factual alignment with source document), Abstractive Relevance (Relev., key information coverage), Extractive Relevance, and Image Selection Relevance (alignment between selected image and summary text).

    Models Abs. Ext. ImgSel
    Consist. Relev. Relev. Relev.
    BART 2.18 2.40 - -
    UniMS 2.32 2.45 2.37 2.44
    w/o Visual Guide 2.26 2.42 2.41 2.44
    w/o LExt\mathcal{L}^{\text{Ext}} 2.22 2.40 - 2.42
    w/o Both 2.19 2.39 - 2.38

    UniMS achieves the highest consistency score (2.32) and relevance score (2.45) for abstractive generation as well as highest image selection relevance (2.44). All improvements over the baselines are statistically significant (p<0.001p < 0.001).

Coverage note — Omitted the qualitative class activation mapping (CAM) attention visualization (Figure 4) and the novel n-gram percentage bar chart (Figure 3), as they provide supplementary qualitative illustrations of the main findings captured in the architectural and quantitative knowls.

References

  1. 1.Anderson, P.; He, X.; Buehler, C.; Teney, D.; Johnson, M.; Gould, S.; and Zhang, L. 2018. Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering. In CVPR, 6077–6086.
  2. 2.Baltrusaitis, T.; Ahuja, C.; and Morency, L. 2019. Multimodal Machine Learning: A Survey and Taxonomy. IEEE Trans. Pattern Anal. Mach. Intell., 41(2): 423–443.
  3. 3.Cao, Z.; Wei, F.; Li, W.; and Li, S. 2017. Faithful to the Original: Fact Aware Neural Abstractive Summarization. CoRR, abs/1711.04434.
  4. 4.Cho, J.; Lei, J.; Tan, H.; and Bansal, M. 2021. Unifying Vision-and-Language Tasks via Text Generation. In ICML, volume 139, 1931–1942.
  5. 5.Devlin, J.; Chang, M.; Lee, K.; and Toutanova, K. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT, 4171–4186.
  6. 6.Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In ICLR.
  7. 7.Dou, Z.; Liu, P.; Hayashi, H.; Jiang, Z.; and Neubig, G. 2021. GSum: A General Framework for Guided Neural Abstractive Summarization. In NAACL-HLT, 4830–4842.
  8. 8.Erkan, G.; and Radev, D. R. 2004. LexRank: Graph-based Lexical Centrality as Salience in Text Summarization. J. Artif. Intell. Res., 22: 457–479.
  9. 9.He, X.; and Deng, L. 2017. Deep Learning for Image-to-Text Generation: A Technical Overview. IEEE Signal Process. Mag., 34(6): 109–116.
  10. 10.Hinton, G. E.; Vinyals, O.; and Dean, J. 2015. Distilling the Knowledge in a Neural Network. CoRR, abs/1503.02531.
  11. 11.Jia, R.; Cao, Y.; Shi, H.; Fang, F.; Yin, P.; and Wang, S. 2021. Flexible Non-Autoregressive Extractive Summarization with Threshold: How to Extract a Non-Fixed Number of Summary Sentences. In AAAI, 13134–13142.
  12. 12.Kim, W.; Son, B.; and Kim, I. 2021. ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision. In ICML, volume 139, 5583–5594.
  13. 13.Kullback, S.; and Leibler, R. A. 1951. On information and sufficiency. The annals of mathematical statistics, 22(1): 79–86.
  14. 14.Lewis, M.; Liu, Y.; Goyal, N.; Ghazvininejad, M.; Mohamed, A.; Levy, O.; Stoyanov, V.; and Zettlemoyer, L. 2020. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In ACL, 7871–7880.
  15. 15.Li, H.; Zhu, J.; Ma, C.; Zhang, J.; and Zong, C. 2017. Multimodal Summarization for Asynchronous Collection of Text, Image, Audio and Video. In EMNLP, 1092–1102.
  16. 16.Lin, C.-Y. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, 74–81.
  17. 17.Liu, Y.; and Lapata, M. 2019. Text Summarization with Pretrained Encoders. In EMNLP-IJCNLP, 3728–3738.
  18. 18.Nallapati, R.; Zhai, F.; and Zhou, B. 2017. SummaRuNNer: A Recurrent Neural Network Based Sequence Model for Extractive Summarization of Documents. In AAAI, 3075–3081.
  19. 19.Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021. Learning Transferable Visual Models From Natural Language Supervision. In ICML, volume 139, 8748–8763.
  20. 20.Ren, S.; He, K.; Girshick, R. B.; and Sun, J. 2017. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. IEEE Trans. Pattern Anal. Mach. Intell., 39(6): 1137–1149.
  21. 21.See, A.; Liu, P. J.; and Manning, C. D. 2017. Get To The Point: Summarization with Pointer-Generator Networks. In ACL, 1073–1083.
  22. 22.Srihari, R. K. 1994. Computational Models for Integrating Linguistic and Visual Information: A Survey. Artif. Intell. Rev., 8(5-6): 349–369.
  23. 23.Wu, Y.; Schuster, M.; Chen, Z.; Le, Q. V.; Norouzi, M.; Macherey, W.; Krikun, M.; Cao, Y.; Gao, Q.; Macherey, K.; Klingner, J.; Shah, A.; Johnson, M.; Liu, X.; Kaiser, L.; Gouws, S.; Kato, Y.; Kudo, T.; Kazawa, H.; Stevens, K.; Kurian, G.; Patil, N.; Wang, W.; Young, C.; Smith, J.; Riesa, J.; Rudnick, A.; Vinyals, O.; Corrado, G.; Hughes, M.; and Dean, J. 2016. Google’s Neural Machine Translation System: Bridging the Gap between Human and Machine Translation. CoRR, abs/1609.08144.
  24. 24.Zhang, Z.; Wang, J.; Sun, Z.; and Yang, Z. 2021. LAMS: A Location-aware Approach for Multimodal Summarization (Student Abstract). In AAAI, 15949–15950.
  25. 25.Zhou, B.; Khosla, A.; Lapedriza, A.; Oliva, A.; and Torralba, A. 2016. Learning Deep Features for Discriminative Localization. In CVPR, 2921–2929.
  26. 26.Zhu, J.; Li, H.; Liu, T.; Zhou, Y.; Zhang, J.; and Zong, C. 2018. MSMO: Multimodal Summarization with Multimodal Output. In EMNLP, 4154–4164.
  27. 27.Zhu, J.; Zhou, Y.; Zhang, J.; Li, H.; Zong, C.; and Li, C. 2020. Multimodal Summarization with Guidance of Multimodal Reference. In AAAI, 9749–9756.

Citation

MLA
Zhang, Z., et al. “UniMS: A Unified Framework for Multimodal Summarization with Knowledge Distillation”. arXiv, 2021, http://arxiv.org/abs/2109.05812v2.
APA
Zhang, Z., Meng, X., Wang, Y., Jiang, X., Liu, Q., & Yang, Z. (2021). UniMS: A Unified Framework for Multimodal Summarization with Knowledge Distillation. arXiv. http://arxiv.org/abs/2109.05812v2
Chicago
Zhang, Z., X. Meng, Y. Wang, X. Jiang, Q. Liu, and Z. Yang. 2021. “UniMS: A Unified Framework for Multimodal Summarization with Knowledge Distillation”. arXiv. http://arxiv.org/abs/2109.05812v2.
Harvard
Zhang, Z. et al. (2021) “UniMS: A Unified Framework for Multimodal Summarization with Knowledge Distillation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2109.05812v2.
Vancouver
1. Zhang Z, Meng X, Wang Y, Jiang X, Liu Q, Yang Z (2021) UniMS: A Unified Framework for Multimodal Summarization with Knowledge Distillation. arXiv

BibTeX

@article{zhang2021unims,
  title = {UniMS: A Unified Framework for Multimodal Summarization with Knowledge Distillation},
  author = {Zhang, Zhengkun and Meng, Xiaojun and Wang, Yasheng and Jiang, Xin and Liu, Qun and Yang, Zhenglu},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2109.05812v2},
  eprint = {2109.05812}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF