A Naturalistic Open Source Movie for Optical Flow Evaluation

Daniel J. ButlerJonas WulffGarrett B. StanleyMichael J. Black

article2012ECCV2,401 citationsKoenderink Prize

Introduces the MPI-Sintel benchmark, a realistic, open-source animated dataset that exposes critical failure modes in state-of-the-art optical flow algorithms by incorporating large displacements, motion blur, and varied rendering passes.

Listen

Optical flow estimation—the tracking of pixel movement across video frames—is vital for computer vision technologies such as autonomous navigation, video editing, and automated surveillance. However, standard hardware sensors cannot directly capture ground truth motion in natural environments, and manual labeling is both inaccurate and impractical. Existing benchmark datasets have become saturated and fail to represent real-world complexity, as they typically feature short sequences, small motions, and simplified visual conditions without blur or realistic atmospheric effects.

The article introduces MPI-Sintel, a new large-scale optical flow benchmark derived from the open-source animated film Sintel, to rigorously evaluate and expose the failure points of current motion-estimation algorithms under realistic and varied visual conditions.

The researchers modified the open-source rendering software Blender to extract dense, pixel-accurate motion vectors across 35 selected video sequences, divided into 1,064 training frames and 564 test frames. To evaluate algorithm sensitivity, sequences were rendered across three progressive complexity levels: an unshaded Albedo pass, an illuminated Clean pass, and a fully shaded Final pass featuring depth-of-field blur, motion blur, and atmospheric fog. The authors validated the dataset's realism by comparing its image luminance, spatial power spectra, and estimated motion statistics to a companion set of 1,473 frames from real-world films and video clips. They also introduced perturbed sequences to discourage unauthorized ground truth reconstruction and benchmarked six established public algorithms across varied motion speeds, occlusion boundaries, and unmatched image regions.

The evaluation revealed several critical findings. First, existing optical flow methods degraded significantly when tested on the richer MPI-Sintel data; algorithms achieving average endpoint errors of less than 0.5 pixels on previous benchmarks experienced roughly a twenty-fold error increase, averaging 9 to 12 pixels. Second, fast motion proved catastrophic across all tested methods, with error rates spiking past 50 to 75 pixels for motions exceeding 40 pixels per frame. Third, unmatched regions—which account for roughly 8.5% of pixels—yielded severe errors averaging over 40 pixels across all models. Fourth, errors increased consistently as pixels approached motion occlusion boundaries. Finally, algorithms performed better on the illuminated Clean pass than on the flat Albedo pass, demonstrating that natural shading and specular cues assist motion estimation rather than hinder it.

These findings demonstrate that contemporary optical flow models are over-optimized for small, simple motions and perform poorly in realistic visual environments. Relying on current algorithms in high-speed or complex visual settings introduces substantial performance and safety risks. Future development should prioritize methods that handle large displacement, temporal consistency across long sequences, and extreme visual effects rather than marginal optimizations on legacy benchmarks.

Researchers and engineers developing vision systems should adopt the MPI-Sintel benchmark and public evaluation website to track algorithm robustness across specific failure modes. Future research directions include expanding the benchmark to higher resolutions, modeling rolling-shutter artifacts, and providing additional evaluation suites for 3D scene flow, surface segmentation, and physical material estimation. While the synthetic nature of the animated movie includes occasional non-physical lighting and object interpenetrations, the statistical alignment with natural film footage confirms that the dataset serves as a highly credible proxy for testing modern vision algorithms.

  • Paper: A Database and Evaluation Methodology for Optical Flow, Simon Baker et al. (2007). Read this predecessor benchmark first to see how its datasets, endpoint-error measures, and evaluation of motion boundaries set the framework that MPI-Sintel expands to more realistic conditions.
Cover for A Naturalistic Open Source Movie for Optical Flow Evaluation

Abstract

Ground truth optical flow is difficult to measure in real scenes with natural motion. As a result, optical flow data sets are restricted in terms of size, complexity, and diversity, making optical flow algorithms difficult to train and test on realistic data. We introduce a new optical flow data set derived from the open source 3D animated short film Sintel. This data set has important features not present in the popular Middlebury flow evaluation: long sequences, large motions, specular reflections, motion blur, defocus blur, and atmospheric effects. Because the graphics data that generated the movie is open source, we are able to render scenes under conditions of varying complexity to evaluate where existing flow algorithms fail. We evaluate several recent optical flow algorithms and find that current highly-ranked methods on the Middlebury evaluation have difficulty with this more complex data set suggesting further research on optical flow estimation is needed. To validate the use of synthetic data, we compare the image- and flow-statistics of Sintel to those of real films and videos and show that they are similar. The data set, metrics, and evaluation website are publicly available.

Table of Contents

  • 1 Introduction
  • 2 Previous Data Sets
  • 3 Design Decisions and Comparison to Middlebury
  • 4 The Sintel Data Set
  • 5 Statistics of Sintel and Natural Movies
  • 6 Analysis
  • 7 Conclusions
  • References

Knowls

  1. Knowl 1 — MPI-Sintel Optical Flow Benchmark Dataset

    experimental setup

    The MPI-Sintel dataset is an optical flow evaluation benchmark derived from the open-source 3D animated film Sintel (produced using the Blender creation suite). It is designed to overcome the scale and complexity limitations of previous benchmarks by providing long sequences, large non-rigid motions (with velocities exceeding 100100 pixels per frame), depth of field blur, motion blur, specular reflections, and atmospheric effects.

    The benchmark consists of 3535 selected clips from the film (where flow is well-defined and physically meaningful), comprising 16281628 total frames with ground truth optical flow at a widescreen resolution of 1024×4361024 \times 436 pixels and 2424 frames per second, stored as 8-bit PNG images:

    • Training set: 2323 sequences containing 10641064 frames with public ground truth forward optical flow fields, occlusion boundary masks, unmatched pixel masks, and invalid pixel maps.
    • Test set: 1212 sequences containing 564564 frames with withheld ground truth optical flow.

    Ground truth motion vectors at every pixel are extracted directly from the 3D scene elements, camera parameters, and object kinematics by modifying Blender's internal motion blur pipeline.

  2. Knowl 2 — Multi-Pass Rendering for Optical Flow Diagnostic Evaluation

    model/method

    To diagnose specific failure modes of optical flow algorithms, the MPI-Sintel dataset renders identical 3D sequences across three distinct rendering passes of increasing visual complexity:

    1. Albedo Pass: Flat, unshaded surfaces with piecewise constant reflectance colors and no illumination or shadow effects. This pass adheres almost strictly to the brightness constancy assumption everywhere except at occlusion boundaries, serving as a baseline to test flow algorithms when brightness constancy holds.
    2. Clean Pass: Adds realistic illumination effects, including smooth surface shading, self-shadowing, ambient cavity darkening, specular highlights, inter-reflections, and mirroring effects.
    3. Final Pass: Full production rendering resembling the released film. Building upon the Clean pass, it incorporates camera depth of field blur, realistic motion blur, atmospheric effects (such as fog and scattering), and color correction.
  3. Knowl 3 — Ground Truth Motion Boundary Extraction

    algorithm

    Thresholding the optical flow gradient alone falsely flags continuous surfaces with steep velocity gradients (such as ground planes undergoing large motions) as occlusion boundaries. Ground truth motion boundaries are defined by combining physical 3D scene boundaries with flow gradient magnitude thresholds.

    Input: 3D scene mesh data, material assignments, depth map z(x,y)z(x, y), forward optical flow field u(x,y)=(u(x,y),v(x,y))\mathbf{u}(x, y) = (u(x, y), v(x, y))
    Output: Binary motion boundary mask MB(x,y)M_B(x, y)
    Detect object boundaries BobjectB_{\text{object}} from graphics element meshes
    Detect material boundaries BmaterialB_{\text{material}} from texture assignments
    Compute normalized depth gradient magnitude Gz(x,y)=∣∇z(x,y)∣z(x,y)G_z(x, y) = \frac{|\nabla z(x, y)|}{z(x, y)}
    Compute depth boundary mask Bdepth=(Gz(x,y)≥τz)B_{\text{depth}} = (G_z(x, y) \ge \tau_z)
    Compute candidate boundary mask Bcandidate=Bobject∪Bmaterial∪BdepthB_{\text{candidate}} = B_{\text{object}} \cup B_{\text{material}} \cup B_{\text{depth}}
    Compute flow gradient magnitude $|
    \nabla \mathbf{u}(x, y)| = \sqrt{\left(\frac{\partial u}{\partial x}\right)^2 + \left(\frac{\partial u}{\partial y}\right)^2 + \left(\frac{\partial v}{\partial x}\right)^2 + \left(\frac{\partial v}{\partial y}\right)^2}$
    Compute flow gradient mask Mflow=(∣∇u(x,y)∣≥2 pixels/frame)M_{\text{flow}} = (|\nabla \mathbf{u}(x, y)| \ge 2\text{ pixels/frame})
    Set MB(x,y)=Bcandidate∩MflowM_B(x, y) = B_{\text{candidate}} \cap M_{\text{flow}}
    return MB(x,y)M_B(x, y)
  4. Knowl 4 — Anti-Cheating Fraud Detection via Geometric and Camera Perturbation

    model/method

    Because the source 3D Blender assets of Sintel are public, an evaluator could potentially reconstruct the ground truth optical flow directly from the 3D project files. To detect fraudulent submissions without closing the benchmark:

    1. Two of the twelve test sequences are randomly altered in scene geometry and camera trajectory.
    2. At every 1010-frame keyframe interval, random 3D spatial offsets in the range [−0.1,+0.1][-0.1, +0.1] meter equivalents are added to the locations of the camera and all stationary (non-animated) objects.
    3. Perturbation coordinates are linearly interpolated between keyframes, creating slightly increased motion and occasional geometry interpenetrations.

    Submitting ground truth flow reconstructed from the original unperturbed Blender files to these perturbed test sequences produces an average endpoint error (EPE) of 2.782.78 pixels (evaluated using Classic+NL-Fast on the Final pass), whereas legitimately running the algorithm on the perturbed sequences yields an EPE of 1.631.63 pixels (compared to 1.481.48 pixels on unperturbed versions). A significant discrepancy between reported errors on perturbed vs. unperturbed sequences serves as a fraud flag.

  5. Knowl 5 — Region-Specific and Motion-Stratified Optical Flow Evaluation Metrics

    definition

    Optical flow accuracy is evaluated using Average Endpoint Error (EPE), defined for an estimated displacement vector (u,v)(u, v) and ground truth vector (uGT,vGT)(u_{\text{GT}}, v_{\text{GT}}) as:

    EPE=(u−uGT)2+(v−vGT)2\text{EPE} = \sqrt{(u - u_{\text{GT}})^2 + (v - v_{\text{GT}})^2}

    To isolate algorithm performance across distinct spatial and dynamic conditions, EPE is partitioned into stratified challenge sub-metrics:

    • Matched vs. Unmatched Pixels: An explicit Unmatched mask identifies pixels visible in frame tt that are not visible in frame t+1t+1 due to occlusion or exiting the frame boundary (accounting for approximately 8.5%8.5\% of pixels in Sintel). Errors are reported separately over Matched, Unmatched, and All pixels.
    • Distance to Motion Boundaries: Using a 2D Euclidean distance transform from the ground truth motion boundaries (excluding unmatched pixels), average EPE is computed across three boundary distance bands: d10d_{10} (d≤10d \le 10 pixels), d10–60d_{10\text{--}60} (10<d<6010 < d < 60 pixels), and d60d_{60} (d≥60d \ge 60 pixels).
    • Speed Regimes: Ground truth flow magnitude r=uGT2+vGT2r = \sqrt{u_{\text{GT}}^2 + v_{\text{GT}}^2} is binned into s10s_{10} (r≤10r \le 10 pixels/frame), s10–40s_{10\text{--}40} (10<r≤4010 < r \le 40 pixels/frame), and s40s_{40} (r>40r > 40 pixels/frame).
  6. Knowl 6 — Optical Flow Algorithm Performance Across MPI-Sintel Challenge Conditions

    data/table

    Average endpoint error (EPE in pixels) on the MPI-Sintel test set across rendering passes, occlusion regions, boundary distances (d10:d≤10 pxd_{10}: d \le 10\text{ px}, d10-60:10<d<60 pxd_{10\text{-}60}: 10 < d < 60\text{ px}, d60:d≥60 pxd_{60}: d \ge 60\text{ px}), and speed regimes (s10:r≤10 ppfs_{10}: r \le 10\text{ ppf}, s10-40:10<r≤40 ppfs_{10\text{-}40}: 10 < r \le 40\text{ ppf}, s40:r≥40 ppfs_{40}: r \ge 40\text{ ppf}):

    Algorithm EPE match unmatch d10d_{10} d10-60d_{10\text{-}60} d60d_{60} s10s_{10} s10-40s_{10\text{-}40} s40s_{40}
    LDOF
    Final 9.15 5.11 42.45 11.13 8.64 8.59 1.49 4.84 57.33
    Clean 7.59 3.49 41.21 9.77 6.98 7.05 0.94 2.91 51.74
    Albedo 7.43 3.18 42.27 9.53 6.64 7.15 1.02 2.57 50.83
    Classic+NL
    Final 9.18 4.87 44.60 11.64 8.88 8.09 1.11 4.50 60.31
    Clean 7.99 3.83 42.23 10.40 7.76 6.84 0.57 2.70 57.41
    Albedo 8.28 4.08 42.89 10.04 8.07 7.49 0.65 2.67 59.54
    HS
    Final 9.64 5.49 43.83 12.27 9.58 8.13 1.88 5.34 58.29
    Clean 8.77 4.59 43.12 11.86 8.91 6.74 1.14 3.86 58.27
    Albedo 9.72 5.33 45.83 12.04 9.73 8.26 1.50 4.51 62.85
    Classic++
    Final 9.99 5.49 47.08 12.76 9.78 8.58 1.40 5.10 64.16
    Clean 8.75 4.33 45.18 11.61 8.61 7.21 0.90 3.30 60.69
    Albedo 9.22 4.64 46.94 11.61 9.02 8.04 1.08 3.33 63.63
    Classic+NL-Fast
    Final 10.12 5.74 46.24 12.37 9.82 9.13 1.09 4.67 67.82
    Clean 9.16 4.81 45.12 11.44 9.02 7.97 0.56 2.82 66.96
    Albedo 9.30 4.89 45.70 11.06 9.15 8.43 0.55 2.82 68.15
    H-L1
    Final 11.95 7.41 49.52 13.89 11.90 10.85 1.16 7.97 74.80
    Clean 12.67 8.07 50.62 14.84 12.91 11.05 0.75 9.98 77.84
    Albedo 12.63 7.98 51.03 14.72 12.90 11.04 0.66 9.67 78.79

    Methods achieving state-of-the-art results on Middlebury (EPE <0.5< 0.5 pixels) exhibit errors roughly 2020 times higher on MPI-Sintel (overall EPE ≈7.4–12.7\approx 7.4\text{--}12.7 pixels). In unmatched occlusion regions, EPE exceeds 4040 pixels across all methods. For fast motions (s40≥40s_{40} \ge 40 ppf), EPE ranges from 50.850.8 to 78.878.8 pixels.

  7. Knowl 7 — Statistical Fidelity of Synthetic Sintel Data Compared to Natural Video

    empirical result

    To evaluate whether synthetic graphics from Sintel serve as a realistic proxy for real-world footage, image and motion statistics were compared across MPI-Sintel, the Middlebury benchmark, and a reference "Lookalikes" dataset comprising 14731473 frames sampled from feature films, TV shows, and web videos clustered into matching semantic scene categories (bamboo, cave, indoor, outdoor, mountain, snowfight):

    • Luminance Distributions: The Kullback-Leibler (KL) divergence of gray-scale pixel intensities relative to the Lookalikes dataset is 0.0580.058 for Sintel, compared to 0.1760.176 for Middlebury.
    • Spatial Power Spectra: The log-log spatial power spectrum slope in the xx-direction (estimated from 2D FFTs on central 436×436436 \times 436 patches) is −2.27-2.27 for Sintel, −2.36-2.36 for Lookalikes, and −2.17-2.17 for Middlebury, matching the characteristic 1/f21/f^2 (slope ≈−2.0\approx -2.0) falloff of natural scenes.
    • Spatial Intensity Derivatives: Horizontal first differences dIdx\frac{dI}{dx} display heavy-tailed log-distributions typical of natural imagery, with a kurtosis of 58.6158.61 for Sintel, 49.9449.94 for Lookalikes, and 23.9323.93 for Middlebury.
    • Flow Statistics: Motion fields computed using Classic+NL-Fast show that Sintel and Lookalikes exhibit comparable high-velocity tails and heavy-tailed motion derivative distributions, both substantially exceeding Middlebury in dynamic complexity.
  8. Knowl 8 — Impact of Illumination, Shading, and Blur Across Render Passes on Optical Flow Estimation

    empirical result

    Systematic evaluation of optical flow methods across rendering passes reveals two primary behaviors:

    1. Illumination Structure Aids Motion Estimation: Although the brightness constancy assumption holds almost everywhere in the Albedo pass, optical flow algorithms (such as Classic+NL, Classic++, and Classic+NL-Fast) generally perform worse on Albedo than on the Clean pass. The Albedo pass produces large, textureless, homogeneous color regions that leave optical flow underconstrained. The smooth shading, self-shadowing, and specular variations in the Clean pass provide stable visual features that disambiguate local correspondences.
    2. Degradation from Atmospheric Blur and Motion Velocity: For most methods, the Final pass is significantly more challenging than the Clean pass due to atmospheric scattering and defocus blur. However, for large-displacement estimation in methods like H-L1, heavy motion blur in the Final pass can act as a natural scale-space smoothing filter, yielding lower high-speed error on the Final pass than on Clean or Albedo.
  9. Knowl 9 — Physical Inconsistencies and Modeling Limitations of the Sintel Dataset

    limitation

    While MPI-Sintel provides ground truth optical flow for complex scenes, several deliberate simplifications and non-physical animation properties limit its applicability for algorithms strictly enforcing real-world physical laws:

    • Non-Physical Lighting and Geometry: Because the source movie prioritized aesthetic appeal over physical simulation, character models occasionally interpenetrate geometry, and characters are lit using local actor-specific illumination (e.g., Sintel's head is lit independently of scene illumination).
    • Exclusion of Optical Transparency: Transparent surfaces require multiple ground truth motion vectors per pixel. To ensure well-defined ground truth, transparent elements were either excluded or modified (e.g., particle-system hair was replaced with opaque polygon strands).
    • Single Motion Vector Representation: Ground truth flow is represented as a single 2D motion vector per pixel, omitting multi-layer flow beneath semi-transparent media or volumetric phenomena.

Coverage note — All substantial contributed material—including dataset construction, rendering passes, boundary and occlusion metrics, anti-cheating methodology, statistical validation, empirical benchmark comparisons, and dataset limitations—is covered across the knowls.

References

  1. 1.Baker, S., Scharstein, D., Lewis, J., Roth, S., Black, M., Szeliski, R.: A database and evaluation methodology for optical flow. IJCV 92, 1–31 (2011)
  2. 2.Roosendaal, T. (Producer): Sintel. Blender Foundation, Durian Open Movie Project (2010), http://www.sintel.org/
  3. 3.http://www.blender.org/
  4. 4.Butler, D., Wulff, J., Stanley, G., Black, M.: MPI-Sintel optical flow benchmark: Supplemental material. MPI-IS-TR-006, MPI for Intelligent Systems (2012)
  5. 5.Roth, S., Black, M.: On the spatial statistics of optical flow. IJCV 74, 33–50 (2007)
  6. 6.Field, D.: Relations between the statistics of natural images and the response properties of cortical cells. J. Opt. Soc. Am. A 4, 2379–2394 (1987)
  7. 7.Simoncelli, E., Olshausen, B.: Natural image statistics and neural representation. Annu. Rev. Neurosci. 24, 1193–1216 (2001)
  8. 8.http://sintel.is.tue.mpg.de
  9. 9.Barron, J., Fleet, D., Beauchemin, S.: Performance of optical flow techniques. IJCV 12, 43–77 (1994)
  10. 10.McCane, B., Novins, K., Crannitch, D., Galvin, B.: On benchmarking optical flow. CVIU 84, 126–143 (2001)
  11. 11.Otte, M., Nagel, H.-H.: Optical Flow Estimation: Advances and Comparisons. In: Eklundh, J.-O. (ed.) ECCV 1994. LNCS, vol. 800, pp. 51–60. Springer, Heidelberg (1994)
  12. 12.Liu, C., Freeman, W., Adelson, E., Weiss, Y.: Human-assisted motion annotation. In: CVPR, pp. 1–8 (2008)
  13. 13.Geiger, A., Lenz, P., Urtasun, R.: Are we ready for autonomous driving? The KITTI vision benchmark suite. In: CVPR, pp. 3354–3361 (2012)
  14. 14.Meister, S., Jaehne, B., Kondermann, D.: An outdoor stereo camera system for the generation of real-world benchmark datasets. Opt. Eng. 51, 021107 (2012)
  15. 15.Sun, D., Roth, S., Lewis, J.P., Black, M.J.: Learning Optical Flow. In: Forsyth, D., Torr, P., Zisserman, A. (eds.) ECCV 2008, Part III. LNCS, vol. 5304, pp. 83–97. Springer, Heidelberg (2008)
  16. 16.Brox, T., Bregler, C., Malik, J.: Large displacement optical flow: Descriptor matching in variational motion estimation. PAMI 33, 500–513 (2009)
  17. 17.Sun, D., Roth, S., Black, M.: Secrets of optical flow estimation and their principles. In: CVPR, pp. 2432–2439 (2010)
  18. 18.Horn, B., Schunck, B.: Determining optical flow. AIJ 16, 185–203 (1981)
  19. 19.Werlberger, M., Trobin, W., Pock, T., Wedel, A., Cremers, D., Bischof, H.: Anisotropic Huber-L1 optical flow. In: BMVC, pp. 1–11 (2009)

Citation

MLA
Butler, D. J., et al. “A Naturalistic Open Source Movie for Optical Flow Evaluation”. Lecture Notes in Computer Science, Springer Berlin Heidelberg, 2012, pp. 611–25, https://doi.org/10.1007/978-3-642-33783-3_44.
APA
Butler, D. J., Wulff, J., Stanley, G. B., & Black, M. J. (2012). A Naturalistic Open Source Movie for Optical Flow Evaluation. In Lecture Notes in Computer Science (pp. 611–625). Springer Berlin Heidelberg. https://doi.org/10.1007/978-3-642-33783-3_44
Chicago
Butler, D. J., J. Wulff, G. B. Stanley, and M. J. Black. 2012. “A Naturalistic Open Source Movie for Optical Flow Evaluation”. In Lecture Notes in Computer Science. Springer Berlin Heidelberg. https://doi.org/10.1007/978-3-642-33783-3_44.
Harvard
Butler, D.J. et al. (2012) “A Naturalistic Open Source Movie for Optical Flow Evaluation”, Lecture Notes in Computer Science. Springer Berlin Heidelberg, pp. 611–625. Available at: https://doi.org/10.1007/978-3-642-33783-3_44.
Vancouver
1. Butler DJ, Wulff J, Stanley GB, Black MJ (2012) A Naturalistic Open Source Movie for Optical Flow Evaluation. In: Lecture Notes in Computer Science. Springer Berlin Heidelberg, pp 611–625

BibTeX

@inbook{Butler_2012, title={A Naturalistic Open Source Movie for Optical Flow Evaluation}, ISBN={9783642337833}, ISSN={1611-3349}, url={http://dx.doi.org/10.1007/978-3-642-33783-3_44}, DOI={10.1007/978-3-642-33783-3_44}, booktitle={Computer Vision – ECCV 2012}, publisher={Springer Berlin Heidelberg}, author={Butler, Daniel J. and Wulff, Jonas and Stanley, Garrett B. and Black, Michael J.}, year={2012}, pages={611–625} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF