UniMS: A Unified Framework for Multimodal Summarization with Knowledge Distillation
Zhengkun ZhangXiaojun MengYasheng WangXin JiangQun LiuZhenglu Yang
Proposes a unified multimodal summarization framework extending BART that jointly handles extractive, abstractive, and image selection tasks while using CLIP-based knowledge distillation to eliminate the need for image captions.
As digital media expands rapidly, users increasingly demand concise pictorial summaries that combine condensed text with the most relevant images. Traditional automated systems typically perform only one type of text summary—either extracting key sentences directly or generating new abstractive phrasing—and rely heavily on high-quality image captions to select pictures. However, real-world multimedia content frequently lacks accurate captions, limiting the practical utility of existing tools.
The article introduces and evaluates UniMS, a unified artificial intelligence framework built on the BART language model. Its primary objective is to simultaneously generate extractive summaries, produce abstractive summaries, and select the most relevant images without needing image captions.
To evaluate this framework, the authors conducted comprehensive experiments on the benchmark MSMO dataset, which contains over 300,000 news articles paired with multiple images. The UniMS approach incorporates linear image patch projections to process visual data efficiently and transfers visual-text matching intelligence from a pretrained vision-language model (CLIP) using knowledge distillation. Furthermore, it incorporates an extractive text reference into the encoder and uses a visually guided decoder to integrate textual and visual signals seamlessly during abstractive text generation.
The experimental and human evaluations established several key findings. First, UniMS established new state-of-the-art performance across all multimodal summarization subtasks, achieving superior text overlap scores and outperforming prior methods in image selection precision (reaching 69.38% precision compared to earlier benchmarks around 59–65%). Second, knowledge distillation proved just as effective as caption-dependent methods, removing the need for manual captions while maintaining high text-image relevance. Third, using lightweight linear patch projections achieved top performance while adding only about 10% parameter overhead, compared to 28–74% extra overhead incurred by complex visual processing backbones. Finally, human evaluation confirmed that UniMS produced significantly more consistent, relevant, and factual summaries than baseline systems.
These results demonstrate that organizations can deploy unified multimodal summarization to process multimedia documents faster and at lower computational costs. By eliminating reliance on image captions and heavy vision backbones, the approach reduces operational overhead and broadens applicability across uncurated multimedia feeds without compromising factual accuracy or relevance.
Decision-makers and technical teams should consider adopting unified encoder-decoder frameworks for multimedia intelligence pipelines, utilizing knowledge distillation to bypass expensive metadata curation like manual captioning. Future technical work should explore pretraining vision-language base models specifically tailored for multimodal summarization tasks to boost performance even further.
Confidence in these findings is high given the robust evaluations on a large standard dataset and rigorous human assessments. A noted limitation is that the model was evaluated primarily on news-style articles with limited visual complexity per story; performance on domain-specific corpora, video inputs, or highly technical documents may require additional validation and fine-tuning.
- Paper: Text Summarization with Pretrained Encoders, Yang Liu et al. (2019). Introduces foundational architectures combining extractive and abstractive summarization within pretrained encoder-decoder frameworks that UniMS adapts into a unified multimodal pipeline.
- Paper: VisualBERT: A Simple and Performant Baseline for Vision and Language, Liunian Harold Li et al. (2019). Establishes a baseline self-attention architecture for jointly modeling visual and textual tokens, providing core visual-linguistic grounding concepts used in UniMS.
- Paper: VL-BERT: Pre-training of Generic Visual-Linguistic Representations, Weijie Su et al. (2019). Demonstrates early methods for generic visual-linguistic pre-training via unified transformer attention, establishing prerequisite multi-modal alignment techniques.
- Paper: SummaRuNNer: A Recurrent Neural Network Based Sequence Model for Extractive Summarization of Documents, Ramesh Nallapati et al. (2016). Provides foundational neural sequence classification principles for extractive document summarization that inform UniMS's extractive text objective.
- Paper: A Neural Attention Model for Abstractive Sentence Summarization, Alexander M. Rush et al. (2015). Pioneers data-driven attention-based neural abstractive summarization, establishing the base sequence-to-sequence generation principles leveraged by UniMS.
- Paper: BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation, Junnan Li et al. (2022). Extends unified vision-language architectures by bootstrapping noisy web data to jointly optimize multimodal understanding and generation.
- Paper: MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models, Deyao Zhu et al. (2024). Builds upon unified visual-text alignment by coupling frozen vision encoders directly with large language models via lightweight linear projection layers.
- Paper: A survey on multimodal large language models, Shukang Yin et al. (2023). Provides a comprehensive architectural survey synthesizing connector designs, distillation, and unified training strategies across modern multimodal large language models.
- Paper: Image as a Foreign Language: BEIT Pretraining for Vision and Vision-Language Tasks, Wenhui Wang et al. (2023). Generalizes unified vision-language modeling by treating image patches directly as foreign language tokens within a shared multiway transformer backbone.
- Paper: Dense Connector for MLLMs, Huanjin Yao et al. (2024). Advances multimodal projection connectors by integrating multi-layer intermediate visual representations into the language model rather than relying solely on final-layer outputs.
- Paper: Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation, Zhiheng Liu et al. (2026). Pushes the unified multimodal framework further by discarding heavy visual encoders entirely in favor of direct patch embedding projections for understanding and generation.
