Built independently by an author, for readers. Read the story and support ChapterPal

keyword

joint moment retrieval

Joint moment retrieval is a video understanding task in artificial intelligence that involves simultaneously identifying the specific time boundaries of an event described by a natural language query and evaluating the importance or saliency of individual video clips within that moment. Rather than treating temporal moment retrieval and highlight detection as separate operations, joint moment retrieval unifies both objectives within a single cross-modal model. This approach processes visual, auditory, and textual signals into a shared representation space, allowing continuous boundary localization and clip-level highlight scoring to inform and reinforce one another. By leveraging the reciprocal relationship between global temporal context and fine-grained clip relevance, joint moment retrieval improves the accuracy of event boundary predictions while pinpointing the most salient moments within extensive video content.

1 item

TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight Detection

TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight Detection

Hao Sun, Mingyao Zhou, Wenjing Chen, Wei Xie

OrganizationsCentral China Normal UniversityHubei University of Technology

Why you should read this

Proposes a DETR-based transformer architecture that improves joint video moment retrieval and highlight detection by directly utilizing task reciprocity through multi-modal feature alignment, visual refinement, and cooperative score exchange across prediction heads.

Video moment retrieval (MR) and highlight detection (HD) based on natural language queries are two highly related tasks, which aim to obtain relevant moments within videos and highlight scores of each video clip. Recently, several methods have been devoted to building DETR-based networks to solve both MR and HD jointly. These methods simply add two separate task heads after multi-modal feature extraction and feature interaction, achieving good performance. Nevertheless, these approaches underutilize the reciprocal relationship between two tasks. In this paper, we propose a task-reciprocal transformer based on DETR (TR-DETR) that focuses on exploring the inherent reciprocity between MR and HD. Specifically, a local-global multi-modal alignment module is first built to align features from diverse modalities into a shared latent space. Subsequently, a visual feature refinement is designed to eliminate query-irrelevant information from visual features for modal interaction. Finally, a task cooperation module is constructed to refine the retrieval pipeline and the highlight score prediction process by utilizing the reciprocity between MR and HD. Comprehensive experiments on QVHighlights, Charades-STA and TVSum datasets demonstrate that TR-DETR outperforms existing state-of-the-art methods. Codes are available at https://github.com/mingyao1120/TR-DETR.

Added

2026-09-26