Built independently by an author, for readers. Read the story and support ChapterPal

keyword

multimodal transformer

A multimodal transformer is a deep learning architecture that adapts the transformer neural network model to simultaneously process, integrate, and analyze information from multiple distinct data modalities, such as text, images, audio, video, and numerical time series. Unlike traditional transformers designed for a single type of data, a multimodal transformer converts heterogeneous inputs into common token representations and projects them into shared or interconnected embedding spaces. It relies on attention mechanisms, including cross-modal and joint self-attention, to capture contextual relationships, semantic alignments, and cross-stream dependencies across unaligned or diverse information streams. By modeling both individual modal characteristics and cross-modal interactions, the architecture provides a unified framework for tasks such as cross-modal retrieval, multimodal sentiment analysis, and joint perceptual reasoning.

7 items

Improving Medical Predictions by Irregular Multimodal Electronic Health Records Modeling

Improving Medical Predictions by Irregular Multimodal Electronic Health Records Modeling

Xinlu Zhang, Shiyang Li, Zhiyu Chen, Xifeng Yan, Linda Ruth Petzold

OrganizationsUniversity of California, Santa Barbara

Why you should read this

Presents a multimodal framework for electronic health records that explicitly captures temporal irregularities across both numerical time series and clinical text through dynamic gating, time-attention mechanisms, and interleaved cross-attention fusion to achieve superior ICU outcome predictions.

Health conditions among patients in intensive care units (ICUs) are monitored via electronic health records (EHRs), composed of numerical time series and lengthy clinical note sequences, both taken at irregular time intervals. Dealing with such irregularity in every modality, and integrating irregularity into multimodal representations to improve medical predictions, is a challenging problem. Our method first addresses irregularity in each single modality by (1) modeling irregular time series by dynamically incorporating handcrafted imputation embeddings into learned interpolation embeddings via a gating mechanism, and (2) casting a series of clinical note representations as multivariate irregular time series and tackling irregularity via a time attention mechanism. We further integrate irregularity in multimodal fusion with an interleaved attention mechanism across temporal steps. To the best of our knowledge, this is the first work to thoroughly model irregularity in multimodalities for improving medical predictions. Our proposed methods for two medical prediction tasks consistently outperforms state-of-the-art (SOTA) baselines in each single modality and multimodal fusion scenarios. Specifically, we observe relative improvements of 6.5%, 3.6%, and 4.3% in F1 for time series, clinical notes, and multimodal fusion, respectively. These results demonstrate the effectiveness of our methods and the importance of considering irregularity in multimodal EHRs.

Added

2026-10-03

Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment Analysis

Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment Analysis

Haoyu Zhang, Yu Wang, Guanghao Yin, Kejun Liu, Yuanyuan Liu, Tianshu Yu

OrganizationsChina University of GeosciencesNanyang Technological UniversityShenzhen Institute of Artificial Intelligence and Robotics for SocietyThe Chinese University of Hong Kong

Why you should read this

Proposes an adaptive language-guided transformer that suppresses conflicting and irrelevant visual and acoustic signals using multi-scale language cues, achieving state-of-the-art multimodal sentiment analysis performance across standard benchmarks like MOSI and MOSEI.

Though Multimodal Sentiment Analysis (MSA) proves effective by utilizing rich information from multiple sources (e.g., language, video, and audio), the potential sentiment-irrelevant and conflicting information across modalities may hinder the performance from being further improved. To alleviate this, we present Adaptive Language-guided Multimodal Transformer (ALMT), which incorporates an Adaptive Hyper-modality Learning (AHL) module to learn an irrelevance/conflict-suppressing representation from visual and audio features under the guidance of language features at different scales. With the obtained hyper-modality representation, the model can obtain a complementary and joint representation through multimodal fusion for effective MSA. In practice, ALMT achieves state-of-the-art performance on several popular datasets (e.g., MOSI, MOSEI and CH-SIMS) and an abundance of ablation demonstrates the validity and necessity of our irrelevance/conflict suppression mechanism.

Added

2026-10-01

Cross-Modal Discrete Representation Learning

Cross-Modal Discrete Representation Learning

Alexander H. Liu, SouYoung Jin, Cheng-I Lai, Andrew Rouditchenko, Aude Oliva, James R. Glass

OrganizationsMassachusetts Institute of Technology

Why you should read this

Presents a self-supervised framework that uses vector quantization and code matching across modalities to learn fine-grained discrete representations, enabling unsupervised concept localization and boosting retrieval performance.

In contrast to recent advances focusing on high-level representation learning across modalities, in this work we present a self-supervised learning framework that is able to learn a representation that captures finer levels of granularity across different modalities such as concepts or events represented by visual objects or spoken words. Our framework relies on a discretized embedding space created via vector quantization that is shared across different modalities. Beyond the shared embedding space, we propose a Cross-Modal Code Matching objective that forces the representations from different views (modalities) to have a similar distribution over the discrete embedding space such that cross-modal objects/actions localization can be performed without direct supervision. We show that the proposed discretized multi-modal fine-grained representation (e.g., pixel/word/frame) can complement high-level summary representations (e.g., video/sentence/waveform) for improved performance on cross-modal retrieval tasks. We also observe that the discretized representation uses individual clusters to represent the same semantic concept across modalities.

Added

2026-09-26

UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition

UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition

Guimin Hu, Ting-En Lin, Yi Zhao, Guangming Lu, Yuchuan Wu, Yongbin Li

OrganizationsHarbin Institute of Technology

Why you should read this

Proposes UniMSE, a generative framework that unifies multimodal sentiment analysis and emotion recognition in conversation by combining label spaces, integrating acoustic and visual signals directly into a T5 backbone, and applying inter-modality contrastive learning.

Multimodal sentiment analysis (MSA) and emotion recognition in conversation (ERC) are key research topics for computers to understand human behaviors. From a psychological perspective, emotions are the expression of affect or feelings during a short period, while sentiments are formed and held for a longer period. However, most existing works study sentiment and emotion separately and do not fully exploit the complementary knowledge behind the two. In this paper, we propose a multimodal sentiment knowledge-sharing framework (UniMSE) that unifies MSA and ERC tasks from features, labels, and models. We perform modality fusion at the syntactic and semantic levels and introduce contrastive learning between modalities and samples to better capture the difference and consistency between sentiments and emotions. Experiments on four public benchmark datasets, MOSI, MOSEI, MELD, and IEMOCAP, demonstrate the effectiveness of the proposed method and achieve consistent improvements compared with state-of-the-art methods.

Added

2026-09-26

Grounded Multimodal Named Entity Recognition on Social Media

Grounded Multimodal Named Entity Recognition on Social Media

Jianfei Yu, Ziyan Li, Jieming Wang, Rui Xia

OrganizationsNanjing University of Science and Technology

Why you should read this

Introduces the task of Grounded Multimodal Named Entity Recognition along with a benchmark Twitter dataset and a hierarchical index generation framework that jointly extracts text entities and locates their corresponding image regions to resolve visual ambiguity in social media posts.

In recent years, Multimodal Named Entity Recognition (MNER) on social media has attracted considerable attention. However, existing MNER studies only extract entity-type pairs in text, which is useless for multimodal knowledge graph construction and insufficient for entity disambiguation. To solve these issues, in this work, we introduce a Grounded Multimodal Named Entity Recognition (GMNER) task. Given a text-image social post, GMNER aims to identify the named entities in text, their entity types, and their bounding box groundings in image (i.e., visual regions). To tackle the GMNER task, we construct a Twitter dataset based on two existing MNER datasets. Moreover, we extend four well-known MNER methods to establish a number of baseline systems and further propose a Hierarchical Index generation framework named H-Index, which generates the entity-type-region triples in a hierarchical manner with a sequence-to-sequence model. Experiment results on our annotated dataset demonstrate the superiority of our H-Index framework over baseline systems on the GMNER task. Our dataset annotation and source code are publicly released at https://github.com/NUSTM/GMNER.

Added

2026-09-26

A Facial Expression-Aware Multimodal Multi-task Learning Framework for Emotion Recognition in Multi-party Conversations

A Facial Expression-Aware Multimodal Multi-task Learning Framework for Emotion Recognition in Multi-party Conversations

Wenjie Zheng, Jianfei Yu, Rui Xia, Shijin Wang

OrganizationsiFLYTEKNanjing University of Science and TechnologyState Key Laboratory of Cognitive Intelligence

Why you should read this

Proposes a two-stage multimodal framework that isolates the true speaker's face sequence from complex multi-party video scenes to accurately guide conversational emotion recognition via multi-task learning.

Multimodal Emotion Recognition in Multi-party Conversations (MERMC) has recently attracted considerable attention. Due to the complexity of visual scenes in multi-party conversations, most previous MERMC studies mainly focus on text and audio modalities while ignoring visual information. Recently, several works proposed to extract face sequences as visual features and have shown the importance of visual information in MERMC. However, given an utterance, the face sequence extracted by previous methods may contain multiple people’s faces, which will inevitably introduce noise to the emotion prediction of the real speaker. To tackle this issue, we propose a two-stage framework named Facial expression-aware Multimodal Multi-Task learning (FacialMMT). Specifically, a pipeline method is first designed to extract the face sequence of the real speaker of each utterance, which consists of multimodal face recognition, unsupervised face clustering, and face matching. With the extracted face sequences, we propose a multimodal facial expression-aware emotion recognition model, which leverages the frame-level facial emotion distributions to help improve utterance-level emotion recognition based on multi-task learning. Experiments demonstrate the effectiveness of the proposed FacialMMT framework on the benchmark MELD dataset. The source code is publicly released at https://github.com/NUSTM/FacialMMT.

Added

2026-09-26

Multimodal Transformer for Unaligned Multimodal Language Sequences

Multimodal Transformer for Unaligned Multimodal Language Sequences

Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J. Zico Kolter, Louis-Philippe Morency, Ruslan Salakhutdinov

OrganizationsBosch Center for AICarnegie Mellon University

Why you should read this

Introduces the Multimodal Transformer to model asynchronous language, audio, and visual streams end-to-end via directional crossmodal attention without requiring explicit word-level alignment preprocessing.

Human language is often multimodal, which comprehends a mixture of natural language, facial gestures, and acoustic behaviors. However, two major challenges in modeling such multimodal human language time-series data exist: 1) inherent data non-alignment due to variable sampling rates for the sequences from each modality; and 2) long-range dependencies between elements across modalities. In this paper, we introduce the Multimodal Transformer (MulT) to generically address the above issues in an end-to-end manner without explicitly aligning the data. At the heart of our model is the directional pairwise crossmodal attention, which attends to interactions between multimodal sequences across distinct time steps and latently adapt streams from one modality to another. Comprehensive experiments on both aligned and non-aligned multimodal time-series show that our model outperforms state-of-the-art methods by a large margin. In addition, empirical analysis suggests that correlated crossmodal signals are able to be captured by the proposed crossmodal attention mechanism in MulT.

Added

2026-09-15