I2SB: Image-to-Image Schrödinger Bridge

Guan-Horng LiuArash VahdatDe-An HuangEvangelos A. TheodorouWeili NieAnima Anandkumar

article2023ICML240 citations

Develops Image-to-Image Schrödinger Bridge (I2^2SB), a simulation-free conditional diffusion framework that directly maps degraded to clean image distributions, outperforming standard diffusion models across restoration tasks without requiring prior knowledge of corruption operators.

Listen

Digital image restoration—such as repairing corrupted, blurry, or low-resolution images—is critical across domains like medical imaging, autonomous systems, data compression, and defense. Standard diffusion-based generative models synthesize clean images by starting from pure random noise and gradually shaping it using the degraded image as a guide. However, starting from unstructured noise ignores the rich structural content already present in the degraded input, resulting in higher computational costs and requiring many sequential processing steps.

The article demonstrates and evaluates the Image-to-Image Schrödinger Bridge (I2SB), a new class of generative diffusion models designed specifically for image-to-image translation. The main objective is to establish an efficient, simulation-free framework that learns direct, nonlinear transitions from degraded images to clean targets without generating content from random noise.

The authors formulated I2SB by reformulating Schrödinger Bridge theory into a tractable structure compatible with standard diffusion architectures. This mathematical approach allows intermediate training states to be computed analytically from clean-degraded image pairs, avoiding the heavy memory and computational overhead of previous Schrödinger Bridge methods. The framework was evaluated on the ImageNet benchmark at 256×256 resolution across multiple restoration tasks, including image inpainting, JPEG artifact removal, deblurring, and 4× super-resolution, and compared against standard conditional diffusion models, traditional Schrödinger Bridge baselines, and specialized inverse solvers.

The evaluation yielded several key findings. First, I2SB outperformed standard conditional diffusion models across most tasks, reducing visual distortion scores (measured by Fréchet Inception Distance) from 8.3 to 4.6 in aggressive JPEG restoration and from 14.8 to 2.8 in bicubic super-resolution. Second, I2SB matched or exceeded the restoration quality of specialized inverse methods without requiring explicit mathematical descriptions of how the images were degraded. Third, the framework drastically improved computational efficiency during generation; for example, on freeform inpainting, I2SB achieved target restoration quality with only 2 to 10 sampling steps, whereas baseline conditional diffusion models required at least 100 steps. Finally, I2SB scaled seamlessly to high-resolution tasks where previous Schrödinger Bridge models were computationally intractable, running significantly faster while cutting memory usage.

These findings indicate that directly modeling transitions between related data distributions provides substantial gains in computational speed and visual fidelity. For organizations deploying computer vision systems, I2SB reduces operational latency and compute costs during inference, enabling near real-time, high-quality image enhancement on standard hardware. Furthermore, eliminating the need to model specific degradation mechanics lowers development complexity across diverse operational settings.

Stakeholders seeking to deploy image restoration pipelines should consider adopting the I2SB framework over conventional noise-to-image diffusion models, particularly where inference budgets or low-latency requirements are strict. Future development should explore combining I2SB with task-specific physical constraints and extending the framework to unsupervised settings where paired training data is unavailable.

A primary limitation of I2SB is its reliance on paired training data (matching degraded and clean images). While paired samples are straightforward to synthesize for most restoration tasks, the approach cannot currently be applied directly to unpaired translation problems without further architectural adaptation. Confidence in the reported performance is high across standard supervised image restoration benchmarks.

Cover for I2SB: Image-to-Image Schrödinger Bridge

Abstract

We propose Image-to-Image Schrödinger Bridge (I2^2SB), a new class of conditional diffusion models that directly learn the nonlinear diffusion processes between two given distributions. These diffusion bridges are particularly useful for image restoration, as the degraded images are structurally informative priors for reconstructing the clean images. I2^2SB belongs to a tractable class of Schrödinger bridge, the nonlinear extension to score-based models, whose marginal distributions can be computed analytically given boundary pairs. This results in a simulation-free framework for nonlinear diffusions, where the I2^2SB training becomes scalable by adopting practical techniques used in standard diffusion models. We validate I2^2SB in solving various image restoration tasks, including inpainting, super-resolution, deblurring, and JPEG restoration on ImageNet 256x256 and show that I2^2SB surpasses standard conditional diffusion models with more interpretable generative processes. Moreover, I2^2SB matches the performance of inverse methods that additionally require the knowledge of the corruption operators. Our work opens up new algorithmic opportunities for developing efficient nonlinear diffusion models on a large scale. scale. Project page and codes: this https URL

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 2.1 Score-based Generative Model (SGM)
  • 2.2 Schrödinger Bridge (SB)
  • 3 Image-to-Image Schrödinger Bridge (I2SB)
  • 3.1 Mathematical Framework
  • 3.2 Algorithmic Design
  • 3.3 Connection to Flow-based Optimal Transport (OT)
  • 3.4 Comparison to Standard Conditional Diffusion Model
  • 4 Related Work
  • 5 Experiment
  • 5.1 Experimental Setup
  • 5.2 Experimental Results
  • 5.3 Discussions
  • 6 Conclusion
  • References
  • A Proof
  • B Introduction to Schrödinger Bridge
  • C Experiment Details
  • C.1 Additional Experimental Setup
  • C.2 Additional Qualitative Results
  • C.3 Additional Discussions

Citation

MLA
Liu, G.-H., et al. “I$^2$SB: Image-to-Image Schrödinger Bridge”. arXiv, 2023, http://arxiv.org/abs/2302.05872v3.
APA
Liu, G.-H., Vahdat, A., Huang, D.-A., Theodorou, E. A., Nie, W., & Anandkumar, A. (2023). I$^2$SB: Image-to-Image Schrödinger Bridge. arXiv. http://arxiv.org/abs/2302.05872v3
Chicago
Liu, G.-H., A. Vahdat, D.-A. Huang, E. A. Theodorou, W. Nie, and A. Anandkumar. 2023. “I$^2$SB: Image-to-Image Schrödinger Bridge”. arXiv. http://arxiv.org/abs/2302.05872v3.
Harvard
Liu, G.-H. et al. (2023) “I$^2$SB: Image-to-Image Schrödinger Bridge”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2302.05872v3.
Vancouver
1. Liu G-H, Vahdat A, Huang D-A, Theodorou EA, Nie W, Anandkumar A (2023) I$^2$SB: Image-to-Image Schrödinger Bridge. arXiv

BibTeX

@article{liu2023image,
  title = {I$^2$SB: Image-to-Image Schrödinger Bridge},
  author = {Liu, Guan-Horng and Vahdat, Arash and Huang, De-An and Theodorou, Evangelos A. and Nie, Weili and Anandkumar, Anima},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2302.05872v3},
  eprint = {2302.05872}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/