Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion

Xunpeng YiHan XuHao ZhangLinfeng TangJiayi Ma

article2024CVPR149 citations

Presents a text-guided framework that incorporates natural language prompts into infrared and visible image fusion, enabling user-interactive control and unified degradation removal without requiring separate restoration models.

Listen

Combining visible and infrared imaging is essential for around-the-clock visual sensing, such as detecting heat signatures at night or capturing clear textures during the day. However, raw source images frequently suffer from environmental degradations, including low lighting, overexposure, blur, and sensor noise. Standard fusion models lack the flexibility to resolve these visual defects dynamically and cannot adapt output characteristics to user-specific operational needs, while multi-stage pipelines that run separate restoration algorithms before fusion are cumbersome, inefficient, and often degrade image quality.

The article demonstrates and evaluates Text-IF, an integrated framework that uses natural language prompts to guide image fusion and degradation removal in a single model. The primary objective is to enable degradation-aware, interactive image processing where plain text commands dynamically adjust the fusion behavior without requiring dedicated preprocessing pipelines or manual parameter tuning.

The researchers designed an architecture that merges a transformer-based vision fusion pipeline with text feature extraction powered by a frozen vision-language model (CLIP). The model was trained on 3,618 image pairs annotated with text prompts and tested on 1,135 image pairs across standard benchmark datasets (MSRS, LLVIP, RoadScene, and MFNet). The evaluation compared Text-IF against leading image fusion frameworks—both with and without upstream specialized restoration tools—across image fidelity, sharpness, naturalness, and downstream object detection accuracy.

The evaluation produced several key findings. First, when operating under standard baseline conditions without text guidance, the model outperformed established competitors across information preservation and fidelity metrics, achieving visual information fidelity scores exceeding 1.0 on key benchmarks compared to competitor scores typically below 0.84. Second, in degraded scenarios involving extreme low light, noise, or overexposure, text guidance enabled the single-model architecture to outperform complex multi-model combinations that joined state-of-the-art restoration tools with fusion models. Third, the resulting fused imagery delivered superior downstream utility, raising object detection mean average precision at threshold 0.50 to 0.941 on low-light benchmarks, surpassing all existing alternatives. Finally, ablation testing confirmed that all components of the composite loss function—covering intensity, structural similarity, color consistency, and edge gradients—are essential to preserve both thermal targets and textural clarity.

These findings indicate that unifying image restoration and cross-modal fusion under language-driven control eliminates the operational overhead of switching among specialized enhancement tools. For organizations deploying autonomous systems, surveillance, or night-vision technologies, this all-in-one approach lowers system complexity, reduces computational latency, and provides non-expert operators with intuitive control over visual outputs.

Organizations evaluating vision systems for complex environments should consider adopting unified text-guided fusion models over rigid multi-stage pipelines. Future initiatives should focus on testing the framework across broader real-time field deployments, expanding supported textual prompts, and integrating the model directly into edge-computing hardware. The reported evidence provides strong confidence in the framework's effectiveness across evaluated benchmarks, though performance boundaries in real-world scenarios with unmodeled compound degradations remain an area for further validation.

Cover for Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion

Abstract

Image fusion aims to combine information from different source images to create a comprehensively representative image. Existing fusion methods are typically helpless in dealing with degradations in low-quality source images and non-interactive to multiple subjective and objective needs. To solve them, we introduce a novel approach that leverages semantic text guidance image fusion model for degradation-aware and interactive image fusion task, termed as Text-IF. It innovatively extends the classical image fusion to the text guided image fusion along with the ability to harmoniously address the degradation and interaction issues during fusion. Through the text semantic encoder and semantic interaction fusion decoder, Text-IF is accessible to the all-in-one infrared and visible image degradation-aware processing and the interactive flexible fusion outcomes. In this way, Text-IF achieves not only multi-modal image fusion, but also multi-modal information fusion. Extensive experiments prove that our proposed text guided image fusion strategy has obvious advantages over SOTA methods in the image fusion performance and degradation treatment. The code is available at https://github.com/XunpengYi/Text-IF.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. The Proposed Method
  • 3.1. Problem Formulation
  • 3.2. Image Fusion Pipeline
  • 3.3. Text Interaction Guidance Architecture
  • 3.4. Loss Functions
  • 4. Experiments
  • 4.1. Implementation Details and Datasets
  • 4.2. Comparison without Text Guidance
  • 4.3. Comparison with Text Guidance
  • 4.4. Performance on High-level Task
  • 4.5. Ablation Experiment
  • 5. Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Text-guided degradation-aware fusion formulation

    model/method

    Text-IF extends conventional infrared–visible image fusion from a fixed mapping to a text-conditioned mapping. Let Ivis∈RH×W×3I_{\mathrm{vis}}\in\mathbb{R}^{H\times W\times 3} be the visible image, Iir∈RH×W×1I_{\mathrm{ir}}\in\mathbb{R}^{H\times W\times 1} the infrared image, TtextT_{\mathrm{text}} a user-provided text description of the fusion task or source-image degradation, and IfI_f the fused RGB image. Conventional fusion is represented as a fixed function of the two images, whereas Text-IF uses a parameterized text-guided function:

    If=Fs-if(Ivis,Iir,Ttext;θn-s).I_f=F_{\mathrm{s\text{-}if}}(I_{\mathrm{vis}},I_{\mathrm{ir}},T_{\mathrm{text}};\theta_{\mathrm{n\text{-}s}}).

    Here, Fs-ifF_{\mathrm{s\text{-}if}} is the semantic text-guided fusion mapping and θn-s\theta_{\mathrm{n\text{-}s}} are its trainable parameters. Conditioning the same network on different text descriptions allows the output to be adapted to user requirements and to restoration of different degradations, rather than producing one predetermined fusion result or requiring a separate restoration model for each degradation.

  2. Knowl 2 — Transformer-based multimodal image fusion pipeline

    model/method

    Text-IF uses separate Transformer/Restormer-based encoders for the visible and infrared inputs, followed by cross-modal attention and a multilevel decoder. The encoders produce feature tensors Fvis=ΦvI(Ivis)F_{\mathrm{vis}}=\Phi^{I}_{v}(I_{\mathrm{vis}}) and Fir=ΦiI(Iir)F_{\mathrm{ir}}=\Phi^{I}_{i}(I_{\mathrm{ir}}), where ΦvI\Phi^{I}_{v} and ΦiI\Phi^{I}_{i} are the visible- and infrared-image encoders.

    For each modality, learned projections produce query, key, and value tensors: (Qv,Kv,Vv)=Φvqkv(Fvis)(Q_v,K_v,V_v)=\Phi^{qkv}_{v}(F_{\mathrm{vis}}) and (Qi,Ki,Vi)=Φiqkv(Fir)(Q_i,K_i,V_i)=\Phi^{qkv}_{i}(F_{\mathrm{ir}}). Cross-attention exchanges the queries between modalities so that each modality is spatially attended using the other modality's keys and values:

    Ff(i)=softmax⁡ ⁣(QvKidk)Vi,Ff(v)=softmax⁡ ⁣(QiKvdk)Vv,F_f^{(i)}=\operatorname{softmax}\!\left(\frac{Q_vK_i}{d_k}\right)V_i, \qquad F_f^{(v)}=\operatorname{softmax}\!\left(\frac{Q_iK_v}{d_k}\right)V_v,

    where dkd_k is the query/key scaling factor and Ff(i)F_f^{(i)} and Ff(v)F_f^{(v)} are the cross-attended feature tensors. Their channel-wise concatenation forms the initial fusion representation Ff0=Concat⁡(Ff(i),Ff(v))F_f^{0}=\operatorname{Concat}(F_f^{(i)},F_f^{(v)}). A self-attention operation further enhances this representation before it enters a decoder composed of repeated Transformer-based decoder blocks and semantic interaction guidance modules. Decoder levels are upsampled to match the encoder resolutions. The workflow diagram on page 3 depicts the two modality-specific encoders, cross-fusion layer, semantic decoder, and final fused image as one integrated pipeline rather than as separate restoration and fusion networks.

  3. Knowl 3 — Frozen CLIP semantics and semantic interaction guidance

    model/method

    Text-IF converts a user description into semantic parameters using a frozen CLIP text encoder and a trainable multilayer perceptron. For a text prompt TtextT_{\mathrm{text}}, the frozen encoder produces Ftext∈RN×LF_{\mathrm{text}}\in\mathbb{R}^{N\times L}:

    Ftext={ΦtI}e(Ttext),F_{\mathrm{text}}=\{\Phi_t^{I}\}_{e}(T_{\mathrm{text}}),

    where NN is the number of text tokens or semantic positions, LL is the text-feature dimension, and the subscript ee indicates frozen pretrained weights. Two MLP output branches transform the text feature into a scale parameter γm\gamma_m and a bias parameter βm\beta_m:

    γm=ΦmI(Ftext),βm=ΦmII(Ftext).\gamma_m=\Phi_m^{I}(F_{\mathrm{text}}), \qquad \beta_m=\Phi_m^{II}(F_{\mathrm{text}}).

    At each semantic interaction guidance stage, these parameters modulate an image-fusion feature Ff(i)F_f^{(i)} through residual scale-and-bias control:

    F^f(i)=(1+γm)⊙Ff(i)+βm,\widehat{F}_f^{(i)}=(1+\gamma_m)\odot F_f^{(i)}+\beta_m,

    where ⊙\odot is element-wise multiplication and F^f(i)\widehat{F}_f^{(i)} is the text-conditioned feature. The modulation is inserted repeatedly into the multistage decoder, allowing semantically different descriptions—such as low light, overexposure, low contrast, or a task specification—to alter the fusion computation while preserving the linguistic consistency of the frozen text encoder.

  4. Knowl 4 — Semantically regulated fusion loss

    equation

    Text-IF trains against manually restored high-quality visible and infrared references, denoted by IvisgI_{\mathrm{vis}}^{g} and IirgI_{\mathrm{ir}}^{g}. For an output image IfI_f of spatial size H×WH\times W, the loss combines four objectives. The element-wise maximum of two images is denoted by max⁡(⋅,⋅)\max(\cdot,\cdot), ∇\nabla is a spatial-gradient operator, and ∥⋅∥1\|\cdot\|_1 is the sum of absolute pixel/channel differences.

    The intensity loss preserves salient high-intensity information:

    Lint=1HW∥If−max⁡(Ivisg,Iirg)∥1.L_{\mathrm{int}}=\frac{1}{HW}\left\|I_f-\max(I_{\mathrm{vis}}^{g},I_{\mathrm{ir}}^{g})\right\|_1.

    The structural-similarity loss uses a text-dependent infrared weight δir(t)\delta_{\mathrm{ir}}(t), where tt is the semantic task/degradation description:

    LSSIM(t)=[1−SSIM⁡(If,Ivisg)]+δir(t)[1−SSIM⁡(If,Iirg)].L_{\mathrm{SSIM}}(t)=\left[1-\operatorname{SSIM}(I_f,I_{\mathrm{vis}}^{g})\right]+\delta_{\mathrm{ir}}(t)\left[1-\operatorname{SSIM}(I_f,I_{\mathrm{ir}}^{g})\right].

    The maximum-gradient loss retains the strongest source-image edges:

    Lgrad=1HW∥∇If−max⁡(∇Ivisg,∇Iirg)∥1.L_{\mathrm{grad}}=\frac{1}{HW}\left\|\nabla I_f-\max(\nabla I_{\mathrm{vis}}^{g},\nabla I_{\mathrm{ir}}^{g})\right\|_1.

    The color-consistency loss compares the chrominance channels after converting RGB images to CbCrCbCr. If FCbCr(⋅)F_{CbCr}(\cdot) extracts the CbCb and CrCr channels, then

    Lcolor=1HW∥FCbCr(If)−FCbCr(Ivisg)∥1.L_{\mathrm{color}}=\frac{1}{HW}\left\|F_{CbCr}(I_f)-F_{CbCr}(I_{\mathrm{vis}}^{g})\right\|_1.

    The complete text-regulated objective is

    Ltotal=αint(t)Lint+αSSIM(t)LSSIM(t)+αgrad(t)Lgrad+αcolor(t)Lcolor,L_{\mathrm{total}}=\alpha_{\mathrm{int}}(t)L_{\mathrm{int}}+\alpha_{\mathrm{SSIM}}(t)L_{\mathrm{SSIM}}(t)+\alpha_{\mathrm{grad}}(t)L_{\mathrm{grad}}+\alpha_{\mathrm{color}}(t)L_{\mathrm{color}},

    where the four α\alpha coefficients are semantic-task-dependent hyperparameters. Thus, text controls not only feature modulation but also the relative balance among saliency, structure, edge, and color objectives.

  5. Knowl 5 — Training and evaluation protocol

    experimental setup

    Text-IF was trained with AdamW at learning rate 0.00010.0001, batch size 1616, and 96×9696\times96 cropped source-image patches. Experiments used PyTorch on an NVIDIA GeForce RTX 3090 GPU. The training data comprised 3,618 infrared–visible image pairs and the test data comprised 1,135 pairs.

    The data were drawn from MSRS, MFNet, RoadScene, and LLVIP. Images with different scenes were selected, manually restored to create high-quality source references, and paired with hundreds of degradation and task descriptions so that semantically varied prompts could be used during training and testing. The degradation cases included low-light and overexposed visible images, as well as low-contrast and noisy infrared images.

    Evaluation used SCD, standard deviation (SD), entropy (EN), visual information fidelity (VIF), quality of gradient-based fusion (QAB/FQ_{AB/F}), CLIP-IQA, NIQE, MUSIQ, BRISQUE, and spatial frequency (SF). Higher values indicate better results for SCD, SD, EN, VIF, QAB/FQ_{AB/F}, CLIP-IQA, MUSIQ, and SF; lower values indicate better results for NIQE and BRISQUE. Comparisons included UMF-CMGR, TarDAL, ReCoNet, MURF, U2Fusion, MetaFusion, and DDFM.

  6. Knowl 6 — Fusion quality without additional text guidance

    data/table

    With a default text input that introduces no additional semantic information, Text-IF was compared directly with seven image-fusion methods on MSRS, LLVIP, and RoadScene. The results show that Text-IF is best on all five reported metrics for MSRS and LLVIP; on RoadScene it has the highest EN, VIF, and QAB/FQ_{AB/F}, while the other methods attain higher SCD or SD. The strongest gains are in information preservation, visual-information fidelity, and edge transfer.

    Method MSRS LLVIP RoadScene
    SCD SD EN VIF QAB/F SCD SD EN VIF QAB/F SCD SD EN VIF QAB/F
    UMF-CMGR 0.981 20.819 5.600 0.430 0.266 1.029 31.501 6.569 0.509 0.352 1.613 36.251 6.973 0.554 0.429
    TarDAL 1.484 35.460 6.347 0.673 0.426 0.817 39.070 5.349 0.330 0.252 1.415 42.609 7.054 0.525 0.391
    ReCoNet 1.191 44.374 3.895 0.438 0.367 1.345 41.234 5.514 0.513 0.364 1.589 37.580 6.822 0.504 0.354
    MURF 0.868 16.431 5.047 0.413 0.327 0.514 21.834 6.051 0.386 0.206 1.576 36.788 6.992 0.484 0.432
    U2Fusion 1.182 23.541 5.246 0.506 0.372 0.757 23.614 5.972 0.552 0.341 1.498 30.969 6.739 0.513 0.467
    MetaFusion 1.486 39.432 6.368 0.726 0.478 1.317 42.446 6.823 0.833 0.493 1.581 50.613 7.223 0.512 0.338
    DDFM 1.550 32.749 5.693 0.622 0.431 1.414 38.346 6.979 0.549 0.220 1.864 44.925 7.226 0.544 0.413
    Text-IF 1.681 44.564 6.789 1.046 0.676 1.591 48.834 7.325 1.011 0.616 1.572 48.962 7.332 0.739 0.578
  7. Knowl 7 — Degradation-aware text-guided fusion results

    data/table

    For degraded inputs, the paper compared one Text-IF model driven by text prompts against pipelines that first applied a degradation-specific restoration model and then one of the seven fusion methods. The restoration models were URetinex for low-light enhancement, AirNet for contrast enhancement, GDID for denoising, and LMPEC for overexposure correction. Text-IF used the same model parameters in all degradation scenarios. The table reports the complete quantitative comparison across five datasets; higher values are preferred for CLIP-IQA, EN, SD, MUSIQ, and SF, while lower values are preferred for NIQE and BRISQUE.

    Method MSRS LLVIP MFNet DN-MSRS RoadScene
    CLIP-IQA EN NIQE EN NIQE MUSIQ SD EN MUSIQ SD EN NIQE SF NIQE BRISQUE
    eir.+UMF-CMGR 0.101 6.316 3.738 7.087 3.891 47.543 23.684 5.414 34.113 21.047 5.645 6.279 11.047 3.792 32.485
    eir.+TarDAL 0.082 5.855 4.750 7.042 3.659 41.735 33.454 6.142 25.120 23.316 5.399 7.353 11.789 3.667 32.436
    eir.+ReCoNet 0.117 7.216 5.769 7.109 4.695 44.187 41.654 5.161 29.299 41.525 4.463 8.631 10.312 4.785 37.775
    eir.+MURF 0.111 5.872 4.199 6.757 4.177 50.589 23.741 5.601 35.626 20.456 5.280 6.549 15.605 3.779 30.594
    eir.+U2Fusion 0.127 6.724 3.997 7.439 3.969 48.481 33.940 5.740 34.255 28.812 4.609 7.185 18.006 4.215 34.577
    eir.+MetaFusion 0.106 7.302 3.584 7.495 3.722 49.620 42.026 6.665 34.762 39.956 6.398 4.337 26.653 3.473 29.500
    eir.+DDFM 0.094 6.723 3.465 7.150 5.184 35.933 30.465 6.480 26.902 27.362 6.120 4.644 10.493 3.717 32.334
    Text-IF 0.132 7.172 3.708 7.391 3.502 48.625 43.933 6.683 35.650 43.448 6.669 4.012 17.766 3.342 29.021

    The degradation benchmark on pages 7–8 shows that Text-IF achieves a strong aggregate result across low-light, low-contrast, noisy, and overexposed inputs, while avoiding the need to select and switch among restoration networks.

  8. Knowl 8 — Improved downstream object detection on fused images

    empirical result

    The authors evaluated whether fusion quality transfers to a high-level vision task by fine-tuning YOLOv8 on infrared–visible source images from LLVIP and evaluating detection on fused images. Text-IF, run with default text and no additional semantic guidance, obtained the best reported detection performance for every metric except that its mAP@0.50:0.95\mathrm{mAP}@0.50{:}0.95 is only slightly above the next-best result.

    Method [email protected] [email protected] [email protected]:0.95
    UMF-CMGR 0.925 0.659 0.599
    TarDAL 0.922 0.646 0.582
    ReCoNet 0.916 0.617 0.568
    MURF 0.926 0.675 0.599
    U2Fusion 0.921 0.655 0.591
    MetaFusion 0.916 0.690 0.590
    DDFM 0.921 0.655 0.592
    Text-IF 0.941 0.676 0.602

    The qualitative detection comparison on page 8 shows that Text-IF detects all visible objects in the displayed scene, whereas competing fusion outputs contain missed detections.

  9. Knowl 9 — Contribution of the four loss components

    data/table

    An LLVIP ablation tested the intensity, structural-similarity, maximum-gradient, and color-consistency losses. The intensity term preserves salient thermal targets, the color term maintains visible-image color, and the gradient term improves texture clarity. The numerical results show that using all four terms gives the highest SCD, SD, and EN among the listed configurations, although omitting color or gradient produces higher individual VIF or QAB/FQ_{AB/F} values.

    LintL_{\mathrm{int}} LSSIML_{\mathrm{SSIM}} LgradL_{\mathrm{grad}} LcolorL_{\mathrm{color}} SCD SD EN VIF QAB/F
    ✓ 1.389 46.147 7.205 0.794 0.552
    ✓ ✓ ✓ 1.481 42.530 7.063 1.020 0.674
    ✓ ✓ ✓ 1.485 47.559 7.182 0.831 0.594
    ✓ ✓ ✓ 1.547 46.798 7.274 0.987 0.688
    ✓ ✓ ✓ ✓ 1.591 48.834 7.325 1.011 0.616

    The qualitative ablation visualization on page 8 associates removal of color consistency with poor color, removal of intensity with insignificant thermal targets, and removal of the gradient term with unclear texture; the complete loss produces the authors' preferred overall fusion appearance.

  10. Knowl 10 — Interactive qualitative behavior across degradation prompts

    empirical result

    The text-guidance comparison uses natural-language descriptions that combine the fusion task with the relevant degradation: low-light visible images in MSRS and LLVIP, low-contrast infrared images in MFNet, noisy infrared images in DN-MSRS, and overexposed visible images in RoadScene. The qualitative comparison on page 7 places Text-IF beside restoration-plus-fusion pipelines for the same source pairs.

    For low-light inputs, Text-IF produces brighter fusion images with more suitable color saturation than pipelines that brighten the visible image before fusion, whose results can remain dark because of the infrared input. For low-contrast infrared or overexposed visible inputs, Text-IF expands the usable dynamic range, preserves visible-image color, and exposes clearer details. For noisy infrared inputs, its fused results contain less residual noise than the GDID-plus-fusion comparisons while retaining prominent thermal targets. These outcomes demonstrate the intended interaction: a user can change the requested restoration/fusion behavior through text without replacing the fusion model or manually selecting a degradation-specific preprocessing network.

Coverage note — No substantial contributed material was omitted; the paper does not state a separate methodological limitation beyond the evaluated datasets, prompts, and degradation cases.

References

  1. 1.Mahmoud Afifi, Konstantinos G Derpanis, Bjorn Ommer, and Michael S Brown. Learning multi-scale photo exposure correction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 9157–9167, 2021. 6
  2. 2.V Aslantas and Emre Bendes. A new image quality metric for image fusion: The sum of the correlations of differences. Aeu-International Journal of Electronics and Communications, 69(12):1890–1896, 2015. 5
  3. 3.Haoyu Chen, Jinjin Gu, Yihao Liu, Salma Abdel Magid, Chao Dong, Qiong Wang, Hanspeter Pfister, and Lei Zhu. Masked image training for generalizable deep image denoising. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 1692–1703, 2023. 6
  4. 4.Ahmet M Eskicioglu and Paul S Fisher. Image quality measures and their performance. IEEE Transactions on Communications, 43(12):2959–2965, 1995. 5
  5. 5.Qishen Ha, Kohei Watanabe, Takumi Karasawa, Yoshitaka Ushiku, and Tatsuya Harada. Mfnet: Towards real-time semantic segmentation for autonomous vehicles with multispectral scenes. In Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 5108–5115, 2017. 5
  6. 6.Yu Han, Yunze Cai, Yin Cao, and Xiaoming Xu. A new image fusion performance metric based on visual information fidelity. Information Fusion, 14(2):127–135, 2013. 5
  7. 7.Zhanbo Huang, Jinyuan Liu, Xin Fan, Risheng Liu, Wei Zhong, and Zhongxuan Luo. Reconet: Recurrent correction network for fast and efficient multi-modality image fusion. In Proceedings of the European Conference on Computer Vision (ECCV), pages 539–555, 2022. 5
  8. 8.Xinyu Jia, Chuang Zhu, Minzhen Li, Wenqi Tang, and Wenli Zhou. Llvip: A visible-infrared paired dataset for low-light vision. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 3496–3504, 2021. 5
  9. 9.Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and improving the image quality of stylegan. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 8110–8119, 2020. 2
  10. 10.Junjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar, and Feng Yang. Musiq: Multi-scale image quality transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 5148–5157, 2021. 5
  11. 11.Gwanghyun Kim, Taesung Kwon, and Jong Chul Ye. Diffusionclip: Text-guided diffusion models for robust image manipulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2426–2435, 2022. 2
  12. 12.Boyun Li, Xiao Liu, Peng Hu, Zhongqin Wu, Jiancheng Lv, and Xi Peng. All-in-one image restoration for unknown corruption. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 17452–17462, 2022. 6
  13. 13.Wentong Liao, Kai Hu, Michael Ying Yang, and Bodo Rosenhahn. Text to image generation with semantic-spatial aware gan. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 18187–18196, 2022. 2
  14. 14.Jinyuan Liu, Xin Fan, Zhanbo Huang, Guanyao Wu, Risheng Liu, Wei Zhong, and Zhongxuan Luo. Target-aware dual adversarial learning and a multi-scenario multi-modality benchmark to fuse infrared and visible for object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5802–5811, 2022. 2, 5
  15. 15.Jinyuan Liu, Zhu Liu, Guanyao Wu, Long Ma, Risheng Liu, Wei Zhong, Zhongxuan Luo, and Xin Fan. Multi-interactive feature learning and a full-time multi-modality benchmark for image fusion and segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 8115–8124, 2023. 1
  16. 16.Nan Liu, Shuang Li, Yilun Du, Antonio Torralba, and Joshua B Tenenbaum. Compositional visual generation with composable diffusion models. In Proceedings of the European Conference on Computer Vision (ECCV), pages 423–439, 2022. 2
  17. 17.Yu Liu, Xun Chen, Rabab K Ward, and Z Jane Wang. Image fusion with convolutional sparse representation. IEEE Signal Processing Letters, 23(12):1882–1886, 2016. 2
  18. 18.Jiayi Ma, Yong Ma, and Chang Li. Infrared and visible image fusion methods and applications: A survey. Information Fusion, 45:153–178, 2019. 1, 5
  19. 19.Jiayi Ma, Han Xu, Junjun Jiang, Xiaoguang Mei, and Xiao-Ping Zhang. Ddcgan: A dual-discriminator conditional generative adversarial network for multi-resolution image fusion. IEEE Transactions on Image Processing, 29:4980–4995, 2020.
  20. 20.Jiayi Ma, Hao Zhang, Zhenfeng Shao, Pengwei Liang, and Han Xu. Ganmcc: A generative adversarial network with multiclassification constraints for infrared and visible image fusion. IEEE Transactions on Instrumentation and Measurement, 70:1–14, 2021. 1
  21. 21.Anish Mittal, Anush Krishna Moorthy, and Alan Conrad Bovik. No-reference image quality assessment in the spatial domain. IEEE Transactions on Image Processing, 21(12):4695–4708, 2012. 5
  22. 22.Anish Mittal, Rajiv Soundararajan, and Alan C Bovik. Making a “completely blind” image quality analyzer. IEEE Signal Processing Letters, 20(3):209–212, 2012. 5
  23. 23.Or Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or, and Dani Lischinski. Styleclip: Text-driven manipulation of stylegan imagery. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 2085–2094, 2021. 2
  24. 24.Vaishnav Potlapalli, Syed Waqas Zamir, Salman Khan, and Fahad Shahbaz Khan. Promptir: Prompting for all-in-one blind image restoration. arXiv preprint arXiv:2306.13090, 2023. 2
  25. 25.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning, pages 8748–8763, 2021. 2
  26. 26.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10684–10695, 2022. 2
  27. 27.Yiming Sun, Bing Cao, Pengfei Zhu, and Qinghua Hu. Detfusion: A detection-driven infrared and visible image fusion network. In Proceedings of the ACM International Conference on Multimedia, pages 4003–4011, 2022. 1
  28. 28.Linfeng Tang, Jiteng Yuan, and Jiayi Ma. Image fusion in the loop of high-level vision tasks: A semantic-aware real-time infrared and visible image fusion network. Information Fusion, 82:28–42, 2022. 1
  29. 29.Linfeng Tang, Jiteng Yuan, Hao Zhang, Xingyu Jiang, and Jiayi Ma. Piafusion: A progressive infrared and visible image fusion network based on illumination aware. Information Fusion, 83:79–92, 2022. 1, 5
  30. 30.Ming Tao, Hao Tang, Fei Wu, Xiao-Yuan Jing, Bing-Kun Bao, and Changsheng Xu. Df-gan: A simple and effective baseline for text-to-image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 16515–16525, 2022. 2
  31. 31.Di Wang, Jinyuan Liu, Xin Fan, and Risheng Liu. Unsupervised misaligned infrared and visible image fusion via cross-modality image generation and registration. In Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI), 2022. 5
  32. 32.Jianyi Wang, Kelvin CK Chan, and Chen Change Loy. Exploring clip for assessing the look and feel of images. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 2555–2563, 2023. 5
  33. 33.Wenhui Wu, Jian Weng, Pingping Zhang, Xu Wang, Wenhan Yang, and Jianmin Jiang. Uretinex-net: Retinex-based deep unfolding network for low-light image enhancement. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5901–5910, 2022. 6
  34. 34.Han Xu, Jiayi Ma, Junjun Jiang, Xiaojie Guo, and Haibin Ling. U2fusion: A unified unsupervised image fusion network. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(1):502–518, 2022. 2, 5
  35. 35.Han Xu, Jiayi Ma, Jiteng Yuan, Zhuliang Le, and Wei Liu. Rfnet: Unsupervised network for mutually reinforcing multi-modal image registration and fusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 19679–19688, 2022. 1
  36. 36.Han Xu, Jiteng Yuan, and Jiayi Ma. Murf: Mutually reinforcing multi-modal image registration and fusion. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(10):12148–12166, 2023. 5
  37. 37.Syed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat, Fahad Shahbaz Khan, and Ming-Hsuan Yang. Restormer: Efficient transformer for high-resolution image restoration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5728–5739, 2022. 3
  38. 38.Hao Zhang and Jiayi Ma. Sdnet: A versatile squeeze-and-decomposition network for real-time image fusion. International Journal of Computer Vision, 129:2761–2785, 2021. 1
  39. 39.Hao Zhang, Han Xu, Xin Tian, Junjun Jiang, and Jiayi Ma. Image fusion meets deep learning: A survey and perspective. Information Fusion, 76:323–336, 2021. 1
  40. 40.Hang Zhao, Orazio Gallo, Iuri Frosio, and Jan Kautz. Loss functions for image restoration with neural networks. IEEE Transactions on Computational Imaging, 3(1):47–57, 2016. 4
  41. 41.Wenda Zhao, Shigeng Xie, Fan Zhao, You He, and Huchuan Lu. Metafusion: Infrared and visible image fusion via meta-feature embedding from object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 13955–13965, 2023. 5
  42. 42.Zixiang Zhao, Haowen Bai, Yuanzhi Zhu, Jiangshe Zhang, Shuang Xu, Yulun Zhang, Kai Zhang, Deyu Meng, Radu Timofte, and Luc Van Gool. Ddfm: denoising diffusion model for multi-modality image fusion. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 8082–8093, 2023. 2, 5
  43. 43.Zhengjie Zhu, Xiaogang Yang, Ruitao Lu, Tong Shen, Xueli Xie, and Tao Zhang. Clf-net: Contrastive learning for infrared and visible image fusion network. IEEE Transactions on Instrumentation and Measurement, 71:1–15, 2022. 1

Citation

MLA
Yi, X., et al. “Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion”. arXiv, 2024, http://arxiv.org/abs/2403.16387v1.
APA
Yi, X., Xu, H., Zhang, H., Tang, L., & Ma, J. (2024). Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion. arXiv. http://arxiv.org/abs/2403.16387v1
Chicago
Yi, X., H. Xu, H. Zhang, L. Tang, and J. Ma. 2024. “Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion”. arXiv. http://arxiv.org/abs/2403.16387v1.
Harvard
Yi, X. et al. (2024) “Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2403.16387v1.
Vancouver
1. Yi X, Xu H, Zhang H, Tang L, Ma J (2024) Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion. arXiv

BibTeX

@article{yi2024text,
  title = {Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion},
  author = {Yi, Xunpeng and Xu, Han and Zhang, Hao and Tang, Linfeng and Ma, Jiayi},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2403.16387v1},
  eprint = {2403.16387}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE