Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion
Xunpeng YiHan XuHao ZhangLinfeng TangJiayi Ma
Presents a text-guided framework that incorporates natural language prompts into infrared and visible image fusion, enabling user-interactive control and unified degradation removal without requiring separate restoration models.
Combining visible and infrared imaging is essential for around-the-clock visual sensing, such as detecting heat signatures at night or capturing clear textures during the day. However, raw source images frequently suffer from environmental degradations, including low lighting, overexposure, blur, and sensor noise. Standard fusion models lack the flexibility to resolve these visual defects dynamically and cannot adapt output characteristics to user-specific operational needs, while multi-stage pipelines that run separate restoration algorithms before fusion are cumbersome, inefficient, and often degrade image quality.
The article demonstrates and evaluates Text-IF, an integrated framework that uses natural language prompts to guide image fusion and degradation removal in a single model. The primary objective is to enable degradation-aware, interactive image processing where plain text commands dynamically adjust the fusion behavior without requiring dedicated preprocessing pipelines or manual parameter tuning.
The researchers designed an architecture that merges a transformer-based vision fusion pipeline with text feature extraction powered by a frozen vision-language model (CLIP). The model was trained on 3,618 image pairs annotated with text prompts and tested on 1,135 image pairs across standard benchmark datasets (MSRS, LLVIP, RoadScene, and MFNet). The evaluation compared Text-IF against leading image fusion frameworks—both with and without upstream specialized restoration tools—across image fidelity, sharpness, naturalness, and downstream object detection accuracy.
The evaluation produced several key findings. First, when operating under standard baseline conditions without text guidance, the model outperformed established competitors across information preservation and fidelity metrics, achieving visual information fidelity scores exceeding 1.0 on key benchmarks compared to competitor scores typically below 0.84. Second, in degraded scenarios involving extreme low light, noise, or overexposure, text guidance enabled the single-model architecture to outperform complex multi-model combinations that joined state-of-the-art restoration tools with fusion models. Third, the resulting fused imagery delivered superior downstream utility, raising object detection mean average precision at threshold 0.50 to 0.941 on low-light benchmarks, surpassing all existing alternatives. Finally, ablation testing confirmed that all components of the composite loss function—covering intensity, structural similarity, color consistency, and edge gradients—are essential to preserve both thermal targets and textural clarity.
These findings indicate that unifying image restoration and cross-modal fusion under language-driven control eliminates the operational overhead of switching among specialized enhancement tools. For organizations deploying autonomous systems, surveillance, or night-vision technologies, this all-in-one approach lowers system complexity, reduces computational latency, and provides non-expert operators with intuitive control over visual outputs.
Organizations evaluating vision systems for complex environments should consider adopting unified text-guided fusion models over rigid multi-stage pipelines. Future initiatives should focus on testing the framework across broader real-time field deployments, expanding supported textual prompts, and integrating the model directly into edge-computing hardware. The reported evidence provides strong confidence in the framework's effectiveness across evaluated benchmarks, though performance boundaries in real-world scenarios with unmodeled compound degradations remain an area for further validation.
- Paper: U2Fusion: A Unified Unsupervised Image Fusion Network, Han Xu et al. (2020). Provides the foundational unified unsupervised framework for multi-modal image fusion that Text-IF builds upon and extends with degradation awareness and text interaction.
- Paper: DenseFuse: A Fusion Approach to Infrared and Visible Images, Hui Li et al. (2018). Establishes classic deep-learning feature extraction and reconstruction pipelines for infrared and visible image fusion, setting baseline paradigms for deep fusion architectures.
- Paper: Edge-Aware Guidance Fusion Network for RGB-Thermal Scene Parsing, Wujie Zhou et al. (2022). Demonstrates guided multimodal fusion between thermal and visible modalities, informing how cross-modal cues can direct selective feature integration.
- Paper: Provable Dynamic Fusion for Low-Quality Multimodal Data, Qingyang Zhang et al. (2023). Offers essential theoretical and algorithmic insights into handling low-quality and degraded inputs in dynamic multimodal fusion.
- Paper: Probing Synergistic High-Order Interaction in Infrared and Visible Image Fusion, Naishan Zheng et al. (2024). Further advances infrared and visible image fusion by investigating synergistic high-order spatial and channel interactions beyond standard fusion mechanisms.
