Video Enhancement with Task-Oriented Flow

Tianfan XueBaian ChenJiajun WuDonglai WeiWilliam T. Freeman

article2017IJCV1,775 citations

Proposes task-oriented flow (TOFlow) and the Vimeo-90K benchmark, demonstrating that learning task-specific motion representations end-to-end outperforms standard optical flow for video frame interpolation, denoising, and super-resolution.

Listen

Video enhancement applications—including frame interpolation, denoising, compression artifact removal, and super-resolution—traditionally rely on estimating standard optical flow to track motion between video frames. However, estimating exact physical motion is computationally expensive and frequently prone to errors caused by lighting changes, motion blur, and occlusions. More critically, physical motion representations are often sub-optimal for restoration tasks because matching visible object boundaries fails to address the inpainting needed for occluded areas or noise removal.

The article evaluates whether learning a task-specific motion representation within an end-to-end framework improves video processing performance. The authors introduce Task-Oriented Flow (TOFlow), an approach that simultaneously estimates motion and enhances video quality, and present Vimeo-90K, a large-scale, high-quality benchmark dataset tailored for low-level video processing tasks.

To demonstrate this, the authors designed a unified neural network composed of three interconnected modules: motion estimation, image transformation via spatial transformer networks, and task-specific image reconstruction. Rather than training motion estimation separately to minimize displacement errors against true motion, the entire network is trained jointly using self-supervision on final output quality. The framework was evaluated on standard public benchmarks as well as the newly curated Vimeo-90K dataset—which contains 89,800 diverse, high-definition video clips (720p or higher)—across frame interpolation, denoising/deblocking, and super-resolution.

The findings show that TOFlow consistently and significantly outperforms both traditional two-step optical flow pipelines and recent deep learning approaches across all evaluated tasks. In frame interpolation, TOFlow achieved superior reconstruction accuracy on the Vimeo benchmark (reaching up to 33.73 dB PSNR compared to 32.02 dB for EpicFlow and 30.10 dB for fixed flow pipelines) and effectively eliminated ghosting artifacts near occlusion boundaries. In video denoising and deblocking, TOFlow demonstrated strong noise resilience and sharper edge retention, outperforming traditional filters like V-BM4D across varying compression and noise levels. In video super-resolution, using only 7 input frames with TOFlow achieved comparable or superior detail recovery relative to legacy methods relying on 30 to 50 frames. Notably, cross-task ablation tests revealed that flow fields trained on one specific task (such as super-resolution) experience sharp performance drops (e.g., around 5 dB) when applied to a different task (such as denoising), confirming that optimal motion fields must be custom-tailored to the target task.

These results imply that engineering pipelines do not need computationally intensive, physically exact optical flow algorithms to achieve state-of-the-art video restoration. By tightly coupling motion estimation directly to image synthesis objectives, systems can achieve higher output fidelity and reduce downstream processing artifacts. The Vimeo-90K dataset also provides a standard, artifact-free benchmark to train and evaluate future deep-learning video enhancement architectures.

Organizations developing video processing or enhancement pipelines should consider adopting joint, end-to-end training of motion and reconstruction modules rather than using off-the-shelf, general-purpose optical flow models. When deploying super-resolution pipelines, using sequences of 5 to 7 frames offers an optimal balance between quality and computational overhead. When adapting this framework, motion modules should be specifically trained on the target application rather than reused across different restoration domains.

The conclusions should be considered in light of certain limitations: the system primarily assumes linear motion between adjacent frames and experiences reduced super-resolution performance when inputs involve complex point-spread blur kernels (showing approximately a 1 to 2 dB drop with box and Gaussian downsampling). Additionally, inference times on tested hardware range from 200 to 400 milliseconds per clip, which may require further optimization for real-time edge deployment.

Cover for Video Enhancement with Task-Oriented Flow

Abstract

Many video enhancement algorithms rely on optical flow to register frames in a video sequence. Precise flow estimation is however intractable; and optical flow itself is often a sub-optimal representation for particular video processing tasks. In this paper, we propose task-oriented flow (TOFlow), a motion representation learned in a self-supervised, task-specific manner. We design a neural network with a trainable motion estimation component and a video processing component, and train them jointly to learn the task-oriented flow. For evaluation, we build Vimeo-90K, a large-scale, high-quality video dataset for low-level video processing. TOFlow outperforms traditional optical flow on standard benchmarks as well as our Vimeo-90K dataset in three video processing tasks: frame interpolation, video denoising/deblocking, and video super-resolution.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Tasks
  • 4 Task-Oriented Flow for Video Processing
  • 4.1 Toy Example
  • 4.2 Flow Estimation Module
  • 4.3 Image Transformation Module
  • 4.4 Image Processing Module
  • 4.5 Training
  • 5 The Vimeo-90K Dataset
  • 6 Evaluation
  • 6.1 Frame Interpolation
  • 6.2 Video Denoising/Deblocking
  • 6.3 Video Super-Resolution
  • 6.4 Flows Learned from Different Tasks
  • 6.5 Accuracy of Retrained Flow
  • 6.6 Different Flow Estimation Network Structure
  • 7 Conclusion
  • References

Knowls

  1. Knowl 1 — Task-Oriented Flow (TOFlow) Video Enhancement Framework

    model/method

    Task-Oriented Flow (TOFlow) is an end-to-end differentiable neural framework for motion-based video processing tasks, including temporal frame interpolation, video denoising/deblocking, and video super-resolution. Rather than treating optical flow estimation as a separate, fixed pre-processing step that estimates true physical object displacement, TOFlow jointly learns motion estimation and video enhancement in a self-supervised, task-specific loop.

    The framework consists of three sequential modules:

    1. Flow Estimation Module: Given NN input video frames (e.g., N=3N=3 for temporal frame interpolation, N=7N=7 for denoising, deblocking, and super-resolution), the module designates a reference frame index (the central frame IrefI_{\text{ref}}) and estimates dense 2D motion fields between each input frame and the reference. For interpolation where the central frame I2I_2 is missing, the module predicts motion fields v21v_{21} and v23v_{23} (from I2I_2 to I1I_1 and I3I_3) directly from the available pair (I1,I3)(I_1, I_3).
    2. Image Transformation Module: Utilizes Spatial Transformer Networks (STNs) featuring differentiable bilinear interpolation layers to warp each non-reference input frame IkI_k (k≠refk \neq \text{ref}) into the spatial coordinates of the reference frame using the predicted motion field vref,kv_{\text{ref}, k}.
    3. Image Processing Module: A task-specific convolutional neural network that takes the set of registered (warped) frames as input (optionally combined with learned occlusion masks) and aggregates them to synthesize the restored target frame I^ref\hat{I}_{\text{ref}}.

    Because all three modules are differentiable, loss gradients from the reconstructed frame output backpropagate through the spatial transformers directly into the flow estimation network. This enables the flow network to discover task-specific motion representations that naturally account for occlusions, noise filtering, and high-frequency edge reconstruction.

  2. Knowl 2 — Vimeo-90K Dataset and Video Processing Benchmarks

    experimental setup

    The Vimeo-90K dataset is a large-scale, high-quality video corpus designed for training and benchmarking low-level video processing algorithms. It comprises 4,278 high-definition videos (720p720\text{p} or higher) collected from Vimeo, selected specifically without inter-frame video compression (such as H.264 temporal compression) so that each frame is compressed independently to eliminate video codec artifacts.

    Using threshold-based shot boundary detection and GIST feature filtering to remove visually redundant scenes, the dataset is partitioned into 89,800 independent video shots. All frames are standardized to a resolution of 448×256448 \times 256 pixels, retaining consecutive frame sequences with average motion magnitudes between 1 and 8 pixels per frame.

    From this dataset, three task-specific benchmarks are defined:

    • Vimeo Interpolation Benchmark: Consists of 73,171 frame triplets ({I1,I2,I3}\{I_1, I_2, I_3\}) extracted from 14,777 video clips. Triplets satisfy three selection criteria: (1) more than 5% of pixels have motion larger than 3 pixels; (2) the L1L_1 photometric difference between the reference frame and the warped frame under optical flow is at most 15 intensity levels (on a [0,255][0, 255] scale); and (3) the average motion difference between neighboring fields (∥v21−v23∥)(\|v_{21} - v_{23}\|) is less than 1 pixel, enforcing linear motion.
    • Vimeo Denoising and Deblocking Benchmark: Consists of 91,701 frame septuplets ({I1,…,I7}\{I_1, \dots, I_7\}) extracted from 38,990 video clips satisfying criteria (1) and (2). For denoising, degraded inputs are generated with additive Gaussian noise (standard deviation σ=0.1\sigma = 0.1 or σ∈{15,25}\sigma \in \{15, 25\}) and mixed noise (Gaussian plus 10% salt-and-pepper noise). For deblocking, clips are compressed using FFmpeg JPEG2000 (format J2k) with quantization factors q∈{20,40,60}q \in \{20, 40, 60\}.
    • Vimeo Super-Resolution Benchmark: Uses the 91,701 frame septuplets downsampled by a spatial factor of 4 from 448×256448 \times 256 to 112×64112 \times 64 via MATLAB imresize bicubic interpolation, as well as box and Gaussian blur kernels.
  3. Knowl 3 — Training Protocol and Loss Optimization for TOFlow

    model/method

    The TOFlow training pipeline employs a phased pre-training scheme followed by end-to-end task-specific joint fine-tuning.

    1. Pre-training the Flow Network:
      • Step 1: The multi-scale motion estimation network is initialized by pre-training on the MPI Sintel dataset using the L1L_1 distance between predicted flow fields and ground truth optical flow vectors.
      • Step 2: For denoising and super-resolution, the flow network is fine-tuned on degraded (noisy or blurred) inputs. For frame interpolation, the network takes frames I1I_1 and I3I_3 as input and is trained to predict the intermediate motion fields v23v_{23} and v21v_{21} by minimizing the L1L_1 error against ground truth flow fields.
    2. Pre-training the Mask Network (Optional, for Interpolation):
      • An occlusion mask network predicting masks m21m_{21} and m23m_{23} from motion fields v21v_{21} and v23v_{23} is pre-trained by minimizing L1L_1 loss against pseudo-ground truth occlusion masks. These masks are computed via forward-backward cycle consistency: for forward flow v12v_{12} and backward flow v21v_{21}, a pixel p∈I1p \in I_1 is labeled occluded if ∥p−v21(p+v12(p))∥2>2\|p - v_{21}(p + v_{12}(p))\|_2 > 2 pixels.
    3. Joint End-to-End Fine-Tuning:
      • All modules (flow estimation, spatial transformation, mask estimation, and image reconstruction) are trained jointly by minimizing the L1L_1 reconstruction loss between the predicted output frame I^\hat{I} and the ground truth frame IgtI_{\text{gt}}: L=∥I^−Igt∥1\mathcal{L} = \|\hat{I} - I_{\text{gt}}\|_1
      • Optimization is performed using the Adam optimizer with a weight decay of 10−410^{-4} for 15 epochs with a batch size of 1. The learning rate is set to 10−410^{-4} for video denoising, deblocking, and super-resolution, and 3×10−43 \times 10^{-4} for temporal frame interpolation. No explicit supervision or loss is applied to the intermediate optical flow fields during joint training.
  4. Knowl 4 — Network Module Architectures of TOFlow

    model/method

    TOFlow implements modular neural network sub-architectures configured for each stage of processing:

    • Flow Estimation Module: Structured as a 4-level spatial pyramid network based on SpyNet. Each pyramid level employs an independent sub-network consisting of 5 convolutional layers with 7×77 \times 7 kernels (zero-padded), batch normalization, and ReLU activations. The layer channel sequence is 32→64→32→16→232 \to 64 \to 32 \to 16 \to 2. The coarsest level receives a zero motion field; subsequent levels take the upsampled flow prediction from the preceding scale concatenated with the image pyramids.
    • Occlusion Mask Module (Interpolation): Structured as a 4-level convolutional pyramid network. Each level contains 5 convolutional layers with 7×77 \times 7 kernels, batch normalization, and ReLU activations (32→64→32→16→232 \to 64 \to 32 \to 16 \to 2 channels). The first level accepts concatenated optical flow fields (4 channels) and outputs 2 mask channels. Higher levels accept concatenated optical flow fields and bilinear-upsampled masks from the preceding level.
    • Image Processing Modules:
      • Frame Interpolation: A residual structure combining an averaging branch and a residual CNN. The averaging branch computes the mean of the two warped frames I21I_{21} and I23I_{23}. The residual CNN processes the warped frames (or masked warped frames I21′,I23′I'_{21}, I'_{23}) through 3 convolutional layers with kernel sizes 9×99 \times 9, 1×11 \times 1, 1×11 \times 1 and channel dimensions 64→64→364 \to 64 \to 3, with ReLU activations after the first two layers. The output is added to the averaged frame: I^2=(I21+I23)/2+ΔI2\hat{I}_2 = (I_{21} + I_{23})/2 + \Delta I_2.
      • Denoising and Deblocking: 3 convolutional layers with kernel sizes 9×99 \times 9, 1×11 \times 1, 1×11 \times 1 and output channel dimensions 64→64→364 \to 64 \to 3 with ReLU activations, operating directly on the 7 registered frames without residual addition.
      • Super-Resolution: Input frames are pre-upsampled by a factor of 4 via bicubic interpolation and warped to the central reference. The processing CNN consists of 4 convolutional layers with kernel sizes 9×99 \times 9, 9×99 \times 9, 1×11 \times 1, 1×11 \times 1 and channel dimensions 64→64→64→364 \to 64 \to 64 \to 3, with ReLU activations on the first three layers.
  5. Knowl 5 — Frame Interpolation Performance on Standard Benchmarks

    empirical result

    TOFlow and its masked variant (TOFlow + Mask) were evaluated against classical two-step optical flow pipelines (EpicFlow, SpyNet, Fixed Flow) and deep learning models (Deep Voxel Flow [DVF], AdaConv, SepConv) on the Vimeo Interpolation Test Set, the DVF Test Set, and the Middlebury Flow Benchmark.

    Methods Vimeo Interp. DVF Dataset
    PSNR (dB) SSIM PSNR (dB) SSIM
    SpyNet 31.95 0.9601 33.60 0.9633
    EpicFlow 32.02 0.9622 33.71 0.9635
    DVF 33.24 0.9627 34.12 0.9631
    AdaConv 32.33 0.9568 — —
    SepConv 33.45 0.9674 34.69 0.9656
    Fixed Flow 29.09 0.9229 31.61 0.9544
    Fixed Flow + Mask 30.10 0.9322 32.23 0.9575
    TOFlow 33.53 0.9668 34.54 0.9666
    TOFlow + Mask 33.73 0.9682 34.58 0.9667
    Methods PMMST DeepFlow SepConv TOFlow TOFlow Mask
    All Regions 5.783 5.965 5.605 5.67 5.49
    Discontinuous 9.545 9.785 8.741 8.82 8.54
    Untextured 2.101 2.045 2.334 2.20 2.17

    On the Vimeo interpolation benchmark, TOFlow + Mask achieves the highest performance (33.73 dB PSNR, 0.9682 SSIM), outperforming SepConv (33.45 dB), DVF (33.24 dB), EpicFlow (32.02 dB), and Fixed Flow (29.09 dB). On the Middlebury flow benchmark (evaluated via sum-of-squared-differences [SSD] square root error, where lower is better), TOFlow + Mask achieves the lowest overall error (5.49) and lowest motion discontinuity error (8.54).

  6. Knowl 6 — Video Denoising and Deblocking Performance

    empirical result

    TOFlow was evaluated on multi-frame video denoising across various noise types (additive Gaussian noise with σ=15\sigma=15 and σ=25\sigma=25, and mixed Gaussian + 10% salt-and-pepper noise) and on video deblocking across JPEG2000 compression quantization factors q∈{20,40,60}q \in \{20, 40, 60\}, comparing against Fixed Flow and V-BM4D.

    Methods Vimeo-Gauss15 Vimeo-Gauss25 Vimeo-Mixed
    PSNR (dB) SSIM PSNR (dB) SSIM PSNR (dB) SSIM
    Fixed Flow 36.25 0.9626 34.74 0.9411 31.85 0.9089
    TOFlow 36.63 0.9628 34.89 0.9518 33.51 0.9395
    Methods Vimeo-BW V-BM4D Dataset
    PSNR (dB) SSIM PSNR (dB) SSIM
    V-BM4D 33.17 0.8756 30.63 0.8759
    TOFlow 36.75 0.9275 30.36 0.8855
    Methods Vimeo-Blocky (q=20q=20) Vimeo-Blocky (q=40q=40) Vimeo-Blocky (q=60q=60)
    PSNR (dB) SSIM PSNR (dB) SSIM PSNR (dB) SSIM
    V-BM4D 35.75 0.9587 33.72 0.9402 32.67 0.9287
    Fixed Flow 36.52 0.9636 34.50 0.9485 33.06 0.9168
    TOFlow 36.92 0.9663 34.97 0.9527 34.02 0.9447

    Across all RGB noise levels, TOFlow outperforms Fixed Flow, exhibiting a 1.66 dB gain under mixed noise (33.51 dB vs. 31.85 dB). On grayscale video denoising (Vimeo-BW), TOFlow outperforms V-BM4D by 3.58 dB PSNR and 0.0519 SSIM. In video deblocking, TOFlow consistently outperforms Fixed Flow and V-BM4D across all compression levels, eliminating block boundary artifacts and ringing around edges.

  7. Knowl 7 — Video Super-Resolution Performance and Frame Count Analysis

    empirical result

    TOFlow was evaluated on 4×4\times multi-frame video super-resolution on the Vimeo-SR and BayesSR datasets, comparing against single-frame bicubic upsampling, DeepSR, BayesSR, SPMC, and Fixed Flow.

    Input Frames Methods Vimeo-SR BayesSR
    PSNR (dB) SSIM PSNR (dB) SSIM
    Full Clip (30–50) DeepSR — — 22.69 0.7746
    BayesSR — — 24.32 0.8486
    1 Frame Bicubic 29.79 0.9036 22.02 0.7203
    7 Frames DeepSR 25.55 0.8498 21.85 0.7535
    BayesSR 24.64 0.8205 21.95 0.7369
    SPMC 32.70 0.9380 21.84 0.7990
    Fixed Flow 31.81 0.9288 22.85 0.7655
    TOFlow 33.08 0.9417 23.54 0.8070
    • Frame Count Impact: Evaluating TOFlow on Vimeo-SR with varying numbers of input frames yields:
      • 3 Frames: 32.66 dB PSNR / 0.9375 SSIM
      • 5 Frames: 33.04 dB PSNR / 0.9415 SSIM
      • 7 Frames: 33.08 dB PSNR / 0.9417 SSIM Performance improves sharply from 3 to 5 frames (+0.38 dB) and levels off between 5 and 7 frames (+0.04 dB).
    • Downsampling Blur Kernel Impact: Evaluating TOFlow across downsampling kernel types yields:
      • Cubic Kernel: 33.08 dB PSNR / 0.9417 SSIM
      • Box Kernel: 32.08 dB PSNR / 0.9372 SSIM
      • Gaussian Kernel (variance 2 px): 31.15 dB PSNR / 0.9314 SSIM The ~1 dB drop per blur kernel occurs because anti-aliasing filtering eliminates high-frequency aliased signals that multi-frame super-resolution relies upon to reconstruct fine details.
  8. Knowl 8 — Task-Specificity and Cross-Task Flow Transfer

    empirical result

    An ablation study evaluating cross-task transferability of TOFlow models demonstrates that the learned motion representations are task-specific and do not generalize across different video enhancement tasks.

    Tasks trained on Tasks evaluated on (PSNR in dB)
    Denoising Deblocking Super-resolution
    Denoising 34.89 36.13 31.30
    Deblocking 25.74 36.92 31.86
    Super-resolution 25.99 31.86 33.08
    Fixed Flow 34.74 36.52 31.81
    EpicFlow 30.43 30.09 28.05

    When a flow network fine-tuned on deblocking or super-resolution is deployed within the denoising pipeline, reconstruction performance collapses by ~9 dB (from 34.89 dB to 25.74 dB or 25.99 dB), introducing prominent noise artifacts. Applying a super-resolution flow network to deblocking causes ringing artifacts. Generic Fixed Flow outperforms cross-task TOFlow models on mismatched tasks (e.g., Fixed Flow achieves 34.74 dB on denoising vs. 25.74 dB from deblocking flow), but is inferior to TOFlow on matching tasks. Qualitative analysis reveals that interpolation flows are globally smooth across occlusions, whereas super-resolution flows produce high-frequency artificial displacements aligned with texture edges.

  9. Knowl 9 — Dissociation Between Optical Flow Accuracy and Video Processing Quality

    empirical result

    Evaluating flow estimation accuracy on the MPI Sintel benchmark reveals that optimizing motion fields directly for downstream video processing tasks degrades their physical optical flow accuracy, while substantially improving video enhancement quality.

    Regions TOFlow (Denoising) TOFlow (Super-res.) Fixed Flow EpicFlow
    All 16.638 16.586 14.120 4.115
    Matched 11.961 11.851 9.322 1.360
    Unmatched 54.724 55.137 53.117 26.595

    Motion accuracy is evaluated using Average End-Point Error (EPE, where lower indicates closer alignment to true physical motion). EpicFlow achieves an EPE of 4.115 across all regions, and Fixed Flow achieves 14.120, both significantly more accurate than TOFlow fine-tuned on denoising (EPE 16.638) or super-resolution (EPE 16.586). However, in downstream evaluations across frame interpolation, denoising, and super-resolution, TOFlow systematically outperforms EpicFlow and Fixed Flow. This establishes that ground-truth physical optical flow is sub-optimal for video restoration tasks because it only aligns visible surfaces, whereas task-oriented flow alters displacement fields to perform neighborhood inpainting, noise averaging, and high-frequency edge reconstruction.

  10. Knowl 10 — Generalization of TOFlow to Alternative Flow Backbones

    empirical result

    To evaluate the architectural generalizability of the TOFlow joint learning paradigm, the SpyNet flow module was replaced with FlowNetC within the TOFlow framework and evaluated across denoising, deblocking, and super-resolution.

    Methods Denoising Deblocking Super-resolution
    PSNR (dB) SSIM PSNR (dB) SSIM PSNR (dB) SSIM
    Fixed Flow 24.685 0.8297 36.028 0.9672 31.834 0.9291
    TOFlow 24.689 0.8374 36.496 0.9700 33.010 0.9411

    Due to memory constraints, FlowNetC estimated optical flow fields at a reduced resolution of 256×192256 \times 192, which were subsequently bilinearly upsampled to the full resolution. Despite the lower absolute metrics caused by this resolution bottleneck, joint task-oriented training (TOFlow) consistently improved reconstruction quality over fixed separate training (Fixed Flow) across all three tasks (e.g., +1.176 dB PSNR in super-resolution, +0.468 dB in deblocking), confirming that the benefits of task-oriented flow optimization are independent of the specific underlying flow network architecture.

Coverage note — None omitted; the knowls cover all core methodological contributions, dataset specifications, architectural details, training procedures, benchmark evaluations across three tasks, and analytical ablations presented in the paper.

References

  1. 1.Abu-El-Haija S, Kothari N, Lee J, Natsev P, Toderici G, Varadarajan B, Vijayanarasimhan S (2016) Youtube-8m: A large-scale video classification benchmark. arXiv:160908675 2
  2. 2.Ahn N, Kang B, Sohn KA (2018) Fast, accurate, and, lightweight super-resolution with cascading residual network. In: European Conference on Computer Vision 3
  3. 3.Aittala M, Durand F (2018) Burst image deblurring using permutation invariant convolutional neural networks. In: European Conference on Computer Vision, pp 731–747 3
  4. 4.Baker S, Scharstein D, Lewis J, Roth S, Black MJ, Szeliski R (2011) A database and evaluation methodology for optical flow. International Journal of Computer Vision 92(1):1–31 1, 3, 4, 7, 8
  5. 5.Brox T, Bruhn A, Papenberg N, Weickert J (2004) High accuracy optical flow estimation based on a theory for warping. In: European Conference on Computer Vision 2
  6. 6.Brox T, Bregler C, Malik J (2009) Large displacement optical flow. In: IEEE Conference on Computer Vision and Pattern Recognition 2, 6
  7. 7.Bulat A, Yang J, Tzimiropoulos G (2018) To learn image super-resolution, use a gan to learn how to do image degradation first. In: European Conference on Computer Vision 3
  8. 8.Butler DJ, Wulff J, Stanley GB, Black MJ (2012) A naturalistic open source movie for optical flow evaluation. In: European Conference on Computer Vision 6, 12
  9. 9.Caballero J, Ledig C, Aitken A, Acosta A, Totz J, Wang Z, Shi W (2017) Real-time video super-resolution with spatio-temporal networks and motion compensation. In: IEEE Conference on Computer Vision and Pattern Recognition 3
  10. 10.Fischer P, Dosovitskiy A, Ilg E, Häusser P, Hazırbas¸ C, Golkov V, van der Smagt P, Cremers D, Brox T (2015) Flownet: Learning optical flow with convolutional networks. In: IEEE International Conference on Computer Vision 2, 3, 12, 13
  11. 11.Ganin Y, Kononenko D, Sungatullina D, Lempitsky V (2016) Deepwarp: Photorealistic image resynthesis for gaze manipulation. In: European Conference on Computer Vision 3
  12. 12.Ghoniem M, Chahir Y, Elmoataz A (2010) Nonlocal video denoising, simplification and inpainting using discrete regularization on graphs. Signal Process 90(8):2445–2455 3
  13. 13.Godard C, Matzen K, Uyttendaele M (2017) Deep burst denoising. In: European Conference on Computer Vision 3
  14. 14.Horn BK, Schunck BG (1981) Determining optical flow. Artif Intell 17(1-3):185–203 2
  15. 15.Huang Y, Wang W, Wang L (2015) Bidirectional recurrent convolutional networks for multi-frame super-resolution. In: Advances in Neural Information Processing Systems 3
  16. 16.Jaderberg M, Simonyan K, Zisserman A, et al (2015) Spatial transformer networks. In: Advances in Neural Information Processing Systems 3, 5
  17. 17.Jiang H, Sun D, Jampani V, Yang MH, Learned-Miller E, Kautz J (2017) Super slomo: High quality estimation of multiple intermediate frames for video interpolation. In: IEEE Conference on Computer Vision and Pattern Recognition 3
  18. 18.Jiang X, Le Pendu M, Guillemot C (2018) Depth estimation with occlusion handling from a sparse set of light field views. In: IEEE International Conference on Image Processing 3
  19. 19.Jo Y, Oh SW, Kang J, Kim SJ (2018) Deep video super-resolution network using dynamic upsampling filters without explicit motion compensation. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 3224–3232 3
  20. 20.Kappeler A, Yoo S, Dai Q, Katsaggelos AK (2016) Video super-resolution with convolutional neural networks. IEEE Transactions on Computational Imaging 2(2):109–122 3
  21. 21.Kingma DP, Ba J (2015) Adam: A method for stochastic optimization. In: International Conference on Learning Representations 6
  22. 22.Li M, Xie Q, Zhao Q, Wei W, Gu S, Tao J, Meng D (2018) Video rain streak removal by multiscale convolutional sparse coding. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 6644–6653 3
  23. 23.Liao R, Tao X, Li R, Ma Z, Jia J (2015) Video super-resolution via deep draft-ensemble learning. In: IEEE Conference on Computer Vision and Pattern Recognition 3, 6, 10, 11
  24. 24.Liu C, Freeman W (2010) A high-quality video denoising algorithm based on reliable motion estimation. In: European Conference on Computer Vision 1, 3
  25. 25.Liu C, Sun D (2011) A bayesian approach to adaptive video super resolution. In: IEEE Conference on Computer Vision and Pattern Recognition 1, 10, 11
  26. 26.Liu C, Sun D (2014) On bayesian adaptive video super resolution. IEEE Transactions on Pattern Analysis and Machine intelligence 36(2):346–360 3, 6
  27. 27.Liu Z, Yeh R, Tang X, Liu Y, Agarwala A (2017) Video frame synthesis using deep voxel flow. In: IEEE International Conference on Computer Vision 2, 3, 5, 7, 8
  28. 28.Lu G, Ouyang W, Xu D, Zhang X, Gao Z, Sun MT (2018) Deep kalman filtering network for video compression artifact reduction. In: European Conference on Computer Vision, pp 568–584 3
  29. 29.Ma Z, Liao R, Tao X, Xu L, Jia J, Wu E (2015) Handling motion blur in multi-frame super-resolution. In: IEEE Conference on Computer Vision and Pattern Recognition 10
  30. 30.Maggioni M, Boracchi G, Foi A, Egiazarian K (2012) Video denoising, deblocking, and enhancement through separable 4-d nonlocal spatiotemporal transforms. IEEE Transactions on Image Processing 21(9):3952–3966 3, 8
  31. 31.Makansi O, Ilg E, Brox T (2017) End-to-end learning of video super-resolution with motion compensation. In: German Conference on Pattern Recognition 3
  32. 32.Mathieu M, Couprie C, LeCun Y (2016) Deep multi-scale video prediction beyond mean square error. In: International Conference on Learning Representations 3
  33. 33.Mémin E, Pérez P (1998) Dense estimation and object-based segmentation of the optical flow with robust techniques. IEEE Transactions on Image Processing 7(5):703–719 2
  34. 34.Mildenhall B, Barron JT, Chen J, Sharlet D, Ng R, Carroll R (2018) Burst denoising with kernel prediction networks. In: IEEE Conference on Computer Vision and Pattern Recognition 3
  35. 35.Nasrollahi K, Moeslund TB (2014) Super-resolution: a comprehensive survey. Machine Vision and Applications 25(6):1423–1468 3
  36. 36.Niklaus S, Liu F (2018) Context-aware synthesis for video frame interpolation. In: IEEE Conference on Computer Vision and Pattern Recognition 3
  37. 37.Niklaus S, Mai L, Liu F (2017a) Video frame interpolation via adaptive convolution. In: IEEE Conference on Computer Vision and Pattern Recognition 3, 7
  38. 38.Niklaus S, Mai L, Liu F (2017b) Video frame interpolation via adaptive separable convolution. In: IEEE International Conference on Computer Vision 7, 8
  39. 39.Oliva A, Torralba A (2001) Modeling the shape of the scene: A holistic representation of the spatial envelope. International Journal of Computer Vision 42(3):145–175 6
  40. 40.Ranjan A, Black MJ (2017) Optical flow estimation using a spatial pyramid network. In: IEEE Conference on Computer Vision and Pattern Recognition 2, 3, 5, 6, 7, 12, 15
  41. 41.Revaud J, Weinzaepfel P, Harchaoui Z, Schmid C (2015) Epicflow: Edge-preserving interpolation of correspondences for optical flow. In: IEEE Conference on Computer Vision and Pattern Recognition 1, 2, 6, 7, 12
  42. 42.Sajjadi MS, Vemulapalli R, Brown M (2018) Frame-recurrent video super-resolution. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 6626–6634 3
  43. 43.Tao X, Gao H, Liao R, Wang J, Jia J (2017) Detail-revealing deep video super-resolution. In: IEEE International Conference on Computer Vision 2, 3, 10, 11
  44. 44.Varghese G, Wang Z (2010) Video denoising based on a spatiotemporal gaussian scale mixture model. IEEE Transactions on Circuits and Systems for Video Technology 20(7):1032–1040 3
  45. 45.Wang TC, Zhu JY, Kalantari NK, Efros AA, Ramamoorthi R (2017) Light field video capture using a learning-based hybrid imaging system. In: SIGGRAPH 3
  46. 46.Wedel A, Cremers D, Pock T, Bischof H (2009) Structure-and motionadaptive regularization for high accuracy optic flow. In: IEEE Conference on Computer Vision and Pattern Recognition 2
  47. 47.Wen B, Li Y, Pfister L, Bresler Y (2017) Joint adaptive sparsity and low-rankness on the fly: an online tensor reconstruction scheme for video denoising. In: IEEE International Conference on Computer Vision (ICCV) 3
  48. 48.Werlberger M, Pock T, Unger M, Bischof H (2011) Optical flow guided tv-l1 video interpolation and restoration. In: International Conference on Energy Minimization Methods in Computer Vision and Pattern Recognition 3
  49. 49.Xu J, Ranftl R, Koltun V (2017) Accurate optical flow via direct cost volume processing. In: IEEE Conference on Computer Vision and Pattern Recognition 2
  50. 50.Xu S, Zhang F, He X, Shen X, Zhang X (2015) Pm-pm: Patchmatch with potts model for object segmentation and stereo matching. IEEE Transactions on Image Processing 24(7):2182–2196 8
  51. 51.Yang R, Xu M, Wang Z, Li T (2018) Multi-frame quality enhancement for compressed video. In: IEEE Conference on Computer Vision and Pattern Recognition, pp 6664–6673 3
  52. 52.Yu JJ, Harley AW, Derpanis KG (2016) Back to basics: Unsupervised learning of optical flow via brightness constancy and motion smoothness. In: European Conference on Computer Vision Workshops 3
  53. 53.Yu Z, Li H, Wang Z, Hu Z, Chen CW (2013) Multi-level video frame interpolation: Exploiting the interaction among different levels. IEEE Transactions on Circuits and Systems for Video Technology 23(7):1235–1248 3
  54. 54.Zhou T, Tulsiani S, Sun W, Malik J, Efros AA (2016) View synthesis by appearance flow. In: European Conference on Computer Vision 3
  55. 55.Zhu X, Wang Y, Dai J, Yuan L, Wei Y (2017) Flow-guided feature aggregation for video object detection. In: IEEE International Conference on Computer Vision 3
  56. 56.Zitnick CL, Kang SB, Uyttendaele M, Winder S, Szeliski R (2004) High-quality video view interpolation using a layered representation. ACM Transactions on Graphics 23(3):600–608 7

Citation

MLA
Xue, T., et al. “Video Enhancement with Task-Oriented Flow”. International Journal of Computer Vision, vol. 127, no. 8, 2019, pp. 1106–25, https://doi.org/10.1007/s11263-018-01144-2.
APA
Xue, T., Chen, B., Wu, J., Wei, D., & Freeman, W. T. (2019). Video Enhancement with Task-Oriented Flow. International Journal of Computer Vision, 127(8), 1106–1125. https://doi.org/10.1007/s11263-018-01144-2
Chicago
Xue, T., B. Chen, J. Wu, D. Wei, and W. T. Freeman. 2019. “Video Enhancement with Task-Oriented Flow”. International Journal of Computer Vision 127 (8): 1106–25. https://doi.org/10.1007/s11263-018-01144-2.
Harvard
Xue, T. et al. (2019) “Video Enhancement with Task-Oriented Flow”, International Journal of Computer Vision, 127(8), pp. 1106–1125. Available at: https://doi.org/10.1007/s11263-018-01144-2.
Vancouver
1. Xue T, Chen B, Wu J, Wei D, Freeman WT (2019) Video Enhancement with Task-Oriented Flow. International Journal of Computer Vision 127:1106–1125

BibTeX

@article{Xue_2019, title={Video Enhancement with Task-Oriented Flow}, volume={127}, ISSN={1573-1405}, url={http://dx.doi.org/10.1007/s11263-018-01144-2}, DOI={10.1007/s11263-018-01144-2}, number={8}, journal={International Journal of Computer Vision}, publisher={Springer Science and Business Media LLC}, author={Xue, Tianfan and Chen, Baian and Wu, Jiajun and Wei, Donglai and Freeman, William T.}, year={2019}, month=Feb, pages={1106–1125} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF