The SYNTHIA Dataset: A Large Collection of Synthetic Images for Semantic Segmentation of Urban Scenes

Germán RosLaura SellartJoanna MaterzynskaDavid VázquezAntonio M. López

article2016CVPR2,482 citations

Introduces a large-scale synthetic dataset of urban driving scenes with automatic pixel-level annotations, demonstrating that joint training on synthetic and real imagery significantly improves deep semantic segmentation accuracy across real-world benchmarks.

Listen

The article addresses the challenge of obtaining large volumes of pixel-level annotated images needed to train deep convolutional neural networks for semantic segmentation in urban driving scenes. Manual annotation is costly and time-consuming, limiting the scale and diversity of real-world datasets available for autonomous driving systems.

The article set out to determine whether realistic synthetic images generated in a virtual city could usefully supplement or replace real images when training such networks, and whether combining the two domains would improve segmentation accuracy on real test data.

The authors created the SYNTHIA dataset containing more than 213,400 synthetic frames rendered from a virtual urban environment, complete with automatic pixel-level labels for 13 classes and varying seasons, lighting, viewpoints, and dynamic objects. They trained two convolutional architectures on SYNTHIA alone and on balanced batches that mixed SYNTHIA with four existing real datasets, then measured performance on held-out real validation images.

Training solely on SYNTHIA produced reasonable accuracy on real test sets, sometimes matching or exceeding models trained on the smaller real datasets alone. Adding SYNTHIA to real training data raised average per-class accuracy by 7 to 18 points across the tested datasets and architectures, with the largest gains for pedestrians, cars, and cyclists. Global pixel accuracy also improved in most cases.

These results indicate that large-scale synthetic data can materially reduce reliance on expensive manual labeling while boosting the reliability of perception systems for autonomous driving. The approach offers a practical route to greater diversity in training data without proportional increases in cost or time.

Further work should focus on refining domain-adaptation techniques, testing at higher image resolutions to better capture small objects such as signs and poles, and extending the virtual environment to additional cities, weather conditions, and sensor configurations. The current experiments were conducted at low resolution, which limits recognition of fine details and may affect the generalizability of the reported gains.

Cover for The SYNTHIA Dataset: A Large Collection of Synthetic Images for Semantic Segmentation of Urban Scenes

Abstract

Vision-based semantic segmentation in urban scenarios is a key functionality for autonomous driving. Recent revolutionary results of deep convolutional neural networks (DCNNs) foreshadow the advent of reliable classifiers to perform such visual tasks. However, DCNNs require learning of many parameters from raw images; thus, having a sufficient amount of diverse images with class annotations is needed. These annotations are obtained via cumbersome, human labour which is particularly challenging for semantic segmentation since pixel-level annotations are required. In this paper, we propose to use a virtual world to automatically generate realistic synthetic images with pixel-level annotations. Then, we address the question of how useful such data can be for semantic segmentationin particular, when using a DCNN paradigm. In order to answer this question we have generated a synthetic collection of diverse urban images, named SYNTHIA, with automatically generated class annotations. We use SYNTHIA in combination with publicly available real-world urban images with manually provided annotations. Then, we conduct experiments with DCNNs that show how the inclusion of SYNTHIA in the training stage significantly improves performance on the semantic segmentation task.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. The SYNTHIA Dataset
  • 3.1. Virtual World Generator
  • 3.2. SYNTHIA­Rand and SYNTHIA­Seqs
  • 4. Semantic Segmentation & Synthetic Images
  • 4.1. Architectures Specification
  • 4.2. Training on Real and Synthetic Data
  • 5. Experimental Evaluation
  • 5.1. Validation Datasets
  • 5.2. Analysis of Results
  • 6. Conclusions
  • References

Knowls

  1. Knowl 1 — SYNTHIA Dataset Overview and Annotation Framework

    definition

    The SYNTHIA (SYNTHetic collection of Imagery and Annotations) dataset is a synthetic benchmark rendered using the Unity platform, designed for training and evaluating visual perception systems—specifically semantic segmentation—in autonomous driving scenarios.

    The dataset provides photo-realistic urban frames featuring diverse architectural environments (street blocks, highways, rural areas, parks), variable weather and seasonal conditions (spring flowers, summer lighting, fall foliage, winter snow), dynamic illumination (sunny, cloudy, dusk, dynamic shadows), and varying camera viewpoints. Each frame contains full-resolution ground truth with pixel-level semantic labels for 13 distinct classes:

    1. Sky
    2. Building
    3. Road
    4. Sidewalk
    5. Fence
    6. Vegetation
    7. Lane-marking
    8. Pole
    9. Car
    10. Traffic sign
    11. Pedestrian
    12. Cyclist
    13. Miscellaneous

    Each rendered frame is also provided with a pixel-aligned depth map.

  2. Knowl 2 — SYNTHIA-Rand and SYNTHIA-Seqs Subsets

    definition

    The SYNTHIA benchmark is organized into two complementary image collections, each tailored for different learning objectives:

    • SYNTHIA-Rand: Contains 13,400 static frames at 960×720960 \times 720 resolution with a 100100^\circ horizontal field of view. The images are captured by a virtual camera positioned at random locations throughout the virtual city with a height constrained to [1.5 m,2.0 m][1.5\text{ m}, 2.0\text{ m}] above the ground and a minimum spatial separation of 10 m10\text{ m} between acquisition stations. At each station, multiple viewpoints and scene permutations (varied dynamic objects, pavement textures, and lighting) are rendered. This subset is designed for training feedforward deep convolutional neural networks.
    • SYNTHIA-Seqs: Contains over 200,000 video frames divided into four sequences (approximately 50,000 frames per sequence, one for each season) recorded from a moving virtual car navigating dynamic traffic. The capture platform comprises two multi-camera pods separated by a baseline of B=0.8 mB = 0.8\text{ m} along the horizontal axis. Each pod includes four monocular cameras (100100^\circ FOV each) facing 9090^\circ apart to provide an omnidirectional 360360^\circ field of view, paired with virtual depth sensors operating across a range of 1.5 m1.5\text{ m} to 50 m50\text{ m}. This subset is designed to enable spatio-temporal modeling and video understanding.
  3. Knowl 3 — Balanced Gradient Contribution (BGC) for Synthetic-to-Real Domain Adaptation

    algorithm

    Balanced Gradient Contribution (BGC) is a domain adaptation training procedure that mixes synthetic and real-world training examples within every mini-batch in a fixed ratio throughout end-to-end network training.

    Rather than relying purely on synthetic training or pre-training on synthetic data followed by sequential fine-tuning on real data, BGC ensures that real-domain statistics guide optimization while synthetic samples act as a rich data regularizer across all iterations.

    Input: Real-world dataset DrealD_{\text{real}}, Synthetic dataset DsynthD_{\text{synth}}, Mini-batch size NN, Real image count NrealN_{\text{real}}, Synthetic image count NsynthN_{\text{synth}} where Nreal+Nsynth=NN_{\text{real}} + N_{\text{synth}} = N, Loss function L\mathcal{L}, Initial network parameters θ\theta
    Output: Trained network parameters θ\theta
    while not converged do
        Sample BrealDrealB_{\text{real}} \sim D_{\text{real}} such that Breal=Nreal|B_{\text{real}}| = N_{\text{real}}
        Sample BsynthDsynthB_{\text{synth}} \sim D_{\text{synth}} such that Bsynth=Nsynth|B_{\text{synth}}| = N_{\text{synth}}
        Form batch B=BrealBsynthB = B_{\text{real}} \cup B_{\text{synth}}
        Compute loss: L(θ)=1N(x,y)BL(f(x;θ),y)L(\theta) = \frac{1}{N} \sum_{(x, y) \in B} \mathcal{L}(f(x; \theta), y)
        Compute gradient: g=θL(θ)g = \nabla_{\theta} L(\theta)
        Update parameter vector θ\theta via Adam optimizer using gg
    end while
    return θ\theta

    In standard practice, each mini-batch of size N=10N = 10 is formed using Nreal=6N_{\text{real}} = 6 real images and Nsynth=4N_{\text{synth}} = 4 synthetic images.

  4. Knowl 4 — Target-Net (T-Net) Architecture for Road Scene Segmentation

    model/method

    Target-Net (T-Net) is a compact convolutional encoder-decoder neural network designed for semantic segmentation in autonomous driving contexts where computational efficiency and low parameter count are required.

    The network consists of three stages:

    1. Contraction (Encoder) Blocks: A cascade of four contraction blocks. Each block comprises a 7×77 \times 7 2D convolution with padding 3 producing 64 channels, Batch Normalization, a ReLU activation, and a 2×22 \times 2 Max-Pooling operation with stride 2. During pooling, the spatial indices of the maximum values are stored for use in the decoder. Contraction layers are initialized from a VGG-F model pre-trained on ILSVRC image classification.
    2. Expansion (Decoder) Blocks: A cascade of four expansion blocks corresponding to the encoder stages. Each block performs an unpooling operation using the stored max-pooling indices, followed by a 7×77 \times 7 2D convolution (64 input and output channels, padding 3), Batch Normalization, and ReLU. Expansion block weights are initialized using the He normal initialization method.
    3. Classifier: A final 1×11 \times 1 convolution mapping 64 channels to LL class prediction maps, followed by a pixel-wise softmax function.
  5. Knowl 5 — Inverse-Frequency Weighted Cross-Entropy Loss and Training Optimization

    model/method

    To address severe class imbalance in urban road scenes—where background categories like road, sky, and building heavily outnumber dynamic classes like pedestrians and cyclists—models are optimized using an inverse-frequency weighted cross-entropy loss.

    For a dataset with CC classes, the loss over all pixels ii is defined as:

    L=ic=1Cwcyi,clog(p^i,c)\mathcal{L} = -\sum_{i} \sum_{c=1}^{C} w_c y_{i,c} \log(\hat{p}_{i,c})

    where yi,c{0,1}y_{i,c} \in \{0, 1\} is the binary indicator that class cc is the true label at pixel ii, p^i,c\hat{p}_{i,c} is the predicted probability for class cc at pixel ii, and the class weight wcw_c is computed as the inverse frequency of class cc in the training set:

    wc=1fcw_c = \frac{1}{f_c}

    where fcf_c is the fraction of total training pixels belonging to class cc.

    Input images are preprocessed by resizing to 180×120180 \times 120 pixels and applying local contrast normalization independently to each color channel to mitigate drastic illumination changes. The network is trained end-to-end using the Adam optimizer.

  6. Knowl 6 — Cross-Domain Generalization of SYNTHIA-Trained Semantic Segmentation Models

    data/table

    When deep convolutional neural networks are trained exclusively on synthetic imagery from the SYNTHIA-Rand dataset (13,400 synthetic images) without exposure to real-world training images, they exhibit strong zero-shot cross-domain generalization on real-world driving benchmarks across 11 semantic classes.

    The table below reports per-class, mean per-class, and global pixel accuracy (%) on the validation splits of CamVid (401 validation images), KITTI (347 validation images), Urban LabelMe (742 validation images), and CBCL StreetScenes (3,347 validation images) for both T-Net and FCN architectures:

    Method Validation Set Sky Building Road Sidewalk Fence Vegetat. Pole Car Sign Pedest. Cyclist Per-Class Global
    T-Net CamVid 66 85 86 67 0 27 55 79 3 75 46 48.9 79.7
    T-Net KITTI 73 78 92 27 0 10 0 64 0 72 14 39.0 61.9
    T-Net U-LabelMe 20 59 92 13 0 22 38 89 1 64 23 38.3 53.4
    T-Net CBCL 74 71 87 25 0 35 21 68 2 42 36 41.8 66.0
    FCN CamVid 78 66 86 72 12 79 17 91 43 78 68 62.5 74.9
    FCN KITTI 56 65 59 26 17 65 32 52 42 73 40 47.1 62.7
    FCN U-LabelMe 31 63 68 40 23 65 39 85 18 71 46 50.0 59.1
    FCN CBCL 71 59 73 32 26 81 40 78 31 63 72 56.9 68.2

    Purely synthetic training generalizes effectively to large structural categories (roads, buildings) and common dynamic objects (cars, pedestrians). In several benchmarks (such as T-Net evaluated on CamVid, U-LabelMe, and CBCL), the model trained solely on synthetic data achieves mean per-class accuracy close to or exceeding that of models trained directly on the target real-world training splits.

  7. Knowl 7 — Semantic Segmentation Improvements via Joint Real and Synthetic Training

    data/table

    Augmenting real-world training datasets with synthetic data from SYNTHIA-Rand using Balanced Gradient Contribution (BGC) yields consistent and substantial improvements in semantic segmentation accuracy compared to training on real data alone.

    The table below compares per-class, mean per-class, and global accuracy (%) on the validation sets of CamVid, KITTI, Urban LabelMe (U-LabelMe), and CBCL StreetScenes when training on real data alone (split T) versus training jointly on real data and SYNTHIA-Rand (A):

    Method Training Set Sky Build. Road Sdwk. Fence Veg. Pole Car Sign Ped. Cyc. Per-Class Global
    T-Net CamVid 99 65 95 52 7 79 5 80 3 26 6 46.3 81.9
    CamVid + SYNTHIA 98 90 91 63 5 83 9 94 0 58 31 56.5 (+10.2) 90.7 (+8.8)
    KITTI 79 83 87 73 0 85 0 69 0 10 0 44.2 80.5
    KITTI + SYNTHIA 89 86 90 58 0 72 0 76 0 66 29 51.6 (+7.4) 80.8 (+0.3)
    U-LabelMe 72 80 75 45 0 62 2 53 0 14 2 36.4 62.4
    U-LabelMe + SYNTHIA 69 77 93 33 0 62 11 77 1 67 24 46.7 (+10.3) 72.1 (+9.7)
    CBCL 62 77 86 41 0 74 5 63 0 7 0 37.9 73.9
    CBCL + SYNTHIA 72 82 90 39 0 58 26 70 5 52 39 48.4 (+10.5) 75.2 (+1.3)
    FCN CamVid 99 65 98 45 27 54 16 77 11 34 25 52.8 78.4
    CamVid + SYNTHIA 97 70 98 66 39 88 41 88 53 75 79 72.1 (+18.3) 83.6 (+5.2)
    KITTI 75 77 77 64 47 84 18 78 5 1 1 51.5 82.3
    KITTI + SYNTHIA 84 81 82 71 60 86 43 83 24 7 32 59.4 (+7.9) 80.8 (-1.5)
    U-LabelMe 93 81 83 57 2 79 41 72 20 71 63 60.1 79.4
    U-LabelMe + SYNTHIA 93 72 81 63 10 76 46 79 49 76 64 64.4 (+4.3) 76.2 (-3.2)
    CBCL 90 77 90 41 2 80 37 84 10 47 31 53.4 79.7
    CBCL + SYNTHIA 82 78 74 56 1 80 20 78 8 77 35 53.5 (+0.2) 75.2 (-4.5)

    Incorporating SYNTHIA-Rand boosts mean per-class accuracy across all architectures and benchmark datasets, with per-class improvements reaching up to +18.3+18.3 percentage points. The most pronounced performance improvements occur in dynamic foreground categories (pedestrians, cyclists, and cars), reflecting the benefit of synthetic data in compensating for sparse object instances in real driving datasets.

  8. Knowl 8 — Resolution Constraints and FCN Global Accuracy Trade-offs under BGC

    limitation

    Two notable empirical constraints characterize the experimental findings:

    1. Low-Resolution Bottleneck: Downsampling input imagery to 180×120180 \times 120 pixels degrades the spatial fidelity of fine-grained and thin scene structures, such as poles, traffic signs, and fences. Consequently, accuracy on these classes remains very low across both real-only and combined training setups.
    2. Global Accuracy Decrements in FCN: Although combining real and synthetic data via Balanced Gradient Contribution (BGC) systematically improves mean per-class accuracy, it leads to slight reductions in global pixel accuracy for FCN on KITTI (1.5%-1.5\%), Urban LabelMe (3.2%-3.2\%), and CBCL (4.5%-4.5\%). This effect is hypothesized to arise from the interaction between BGC and FCN's multi-resolution skip-connection upsampling scheme, which fuses feature representations across early and late layers.

Coverage note — No substantial contributed material was omitted. Introductory remarks on autonomous driving trends and related work on indoor/CAD-based synthetic data were excluded per extraction rules.

References

  1. 1.M. Aubry, D. Maturana, A. Efros, B. Russell, and J. Sivic. Seeing 3d chairs: exemplar part-based 2d-3d alignment using a large dataset of cad models. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2014.
  2. 2.V. Badrinarayanan, A. Handa, and R. Cipolla. SegNet: A deep convolutional encoder-decoder architecture for robust semantic pixel-wise labelling. arXiv preprint abs/1505.07293, 2015.
  3. 3.S. Bileschi. CBCL StreetScenes challenge framework, 2007.
  4. 4.G. J. Brostow, J. Fauqueur, and R. Cipolla. Semantic object classes in video: A high-definition ground truth database. Pattern Recognition Letters, 2009.
  5. 5.G. J. Brostow, J. Shotton, and R. Cipolla. Segmentation and recognition using structure from motion point clouds. In Eur. Conf. on Computer Vision (ECCV), 2008.
  6. 6.P. P. Busto, J. Liebelt, and J. Gall. Adaptation of synthetic data for coarse-to-fine viewpoint refinement. In British Machine Vision Conf. (BMVC), 2015.
  7. 7.K. Chatfield, K. Simonyan, A. Vedaldi, and A. Zisserman. Return of the devil in the details: Delving deep into convolutional networks. In British Machine Vision Conf. (BMVC), 2014.
  8. 8.M. Cordts, M. Omran, S. Ramos, T. Scharw¨achter, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele. The cityscapes dataset. In CVPR, Workshop, 2015.
  9. 9.A. Geiger, P. Lenz, C. Stiller, and R. Urtasun. Vision meets Robotics: The KITTI Dataset. Intl. J. of Robotics Research, 2013.
  10. 10.R. Girshick, J. Donahue, T. Darrell, and J. Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2014.
  11. 11.A. Handa, V. Patraucean, V. Badrinarayanan, S. Stent, and R. Cipolla. Synthcam3d: Semantic understanding with synthetic indoor scenes. arXiv preprint abs/1505.00171, 2015.
  12. 12.H. Hattori, V. Naresh Boddeti, K. M. Kitani, and T. Kanade. Learning scene-specific pedestrian detectors without real data. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2015.
  13. 13.K. He, X. Zhang, S. Ren, and J. Sun. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. 2015.
  14. 14.B. Kaneva, A. Torralba, and W. T. Freeman. Evaluating image feaures using a photorealistic virtual world. In Intl. Conf. on Computer Vision (ICCV), 2011.
  15. 15.A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei. Large-scale video classification with convolutional neural networks. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2014.
  16. 16.D. P. Kingma and J. Ba. Adam: A method for stochastic optimization. In International Conference on Learning Representations (ICLR), 2015.
  17. 17.A. Krizhevsky, I. Sustkever, and G. E. Hinton. ImageNet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems, 2012.
  18. 18.A. Kundu, Y. Li, F. Dellaert, F. Li, and J. M. Rehg. Joint semantic segmentation and 3D reconstruction from monocular video. In Eur. Conf. on Computer Vision (ECCV), 2014.
  19. 19.T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Doll´ar, and C. L. Zitnick. Microsoft COCO: Common Objects in Context. In Eur. Conf. on Computer Vision (ECCV), 2014.
  20. 20.J. Long, E. Shelhamer, and T. Darrell. Fully convolutional networks for semantic segmentation. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2015.
  21. 21.J. Marin, D. Vazquez, D. Geronimo, and A. Lopez. Learning appearance in virtual scenarios for pedestrian detection. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2010.
  22. 22.R. Mottaghi, X. Chen, X. Liu, N.-G. Cho, S.-W. Lee, S. Fidler, R. Urtasun, and A. Yuille. The role of context for object detection and semantic segmentation in the wild. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2014.
  23. 23.P. K. Nathan Silberman, Derek Hoiem and R. Fergus. Indoor segmentation and support inference from rgbd images. In Eur. Conf. on Computer Vision (ECCV), 2012.
  24. 24.H. Noh, S. Hong, and B. Han. Learning deconvolution network for semantic segmentation. Intl. Conf. on Computer Vision (ICCV), 2015.
  25. 25.P. Panareda, J. Liebelt, and J. Gall. Adaptation of synthetic data for coarse-to-fine viewpoint refinement. In British Machine Vision Conf. (BMVC), 2015.
  26. 26.J. Papon and M. Schoeler. Semantic pose using deep networks trained on synthetic RGB-D. In Intl. Conf. on Computer Vision (ICCV), 2015.
  27. 27.X. Peng, B. Sun, K. Ali, and K. Saenko. Learning deep object detectors from 3D models. In Intl. Conf. on Computer Vision (ICCV), 2015.
  28. 28.L. Pishchulin, A. Jain, M. Andriluka, T. Thormahlen, and B. Schiele. Articulated people detection and pose estimation: reshaping the future. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2012.
  29. 29.G. Ros, S. Ramos, M. Granados, A. Bakhtiary, D. V´azquez, and A. M. L ´opez. Vision-based offline-online perception paradigm for autonomous driving. In Winter Conference on Applications of Computer Vision (WACV), 2015.
  30. 30.G. Ros, S. Stent, P. F. Alcantarilla, and T. Watanabe. Training constrained deconvolutional networks for road scene semantic segmentation. arXiv preprint abs/1604.01545, 2016.
  31. 31.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei. ImageNet large scale visual recognition challenge. Intl. J. of Computer Vision, 2015.
  32. 32.B. C. Russell, A. Torralba, K. P. Murphy, and W. T. Freeman. LabelMe: a database and web-based tool for image annotation. Intl. J. of Computer Vision, 2008.
  33. 33.T. Scharwchter, M. Enzweiler, U. Franke, and S. Roth. Efficient multi-cue scene segmentation. In Pattern Recognition. 2013.
  34. 34.J. Shotton, A. Fitzgibbon, M. Cook, T. Sharp, M. Finocchio, R. Moore, A. Kipman, and A. Blake. Real-time human pose recognition in parts from a single depth image. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2011.
  35. 35.K. Simonyan and A. Zisserman. Two-stream convolutional networks for action recognition in videos. In arXiv preprint abs/1406.2199, 2014.
  36. 36.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. In International Conference on Learning Representations (ICLR), 2015.
  37. 37.B. Sun and K. Saenko. From virtual to reality: Fast adaptation of virtual object detectors to real domains. In British Machine Vision Conf. (BMVC), 2014.
  38. 38.U. Technologies. Unity Development Platform.
  39. 39.T.L Berg, A. Sorokin, G. Wang, D.A. Forsyth, D. Hoeiem, I. Endres, and A. Farhadi. It’s all about the data. Proceedings of the IEEE, 2010.
  40. 40.J. Tompson, A. Jain, Y. LeCun, and C. Bregler. Joint training of a convolutional network and a graphical model for human pose estimation. In Advances in Neural Information Processing Systems, 2014.
  41. 41.A. Torralba and A. Efros. Unbiased look at dataset bias. In IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2011.
  42. 42.D. V´azquez, A. L ´opez, J. Mar´ın, D. Ponsa, and D. Ger´onimo. Virtual and real world adaptation for pedestrian detection. IEEE Trans. Pattern Anal. Machine Intell., 2014.
  43. 43.L. Woensel and G. Archer. Ten technologies which could change our lives. Technical report, EPRS - European Parlimentary Research Service, January 2015.
  44. 44.J. Xu, S. Ramos, D. Vazquez, and A. Lopez. Domain adaptation of deformable part-based models. IEEE Trans. Pattern Anal. Machine Intell., 2014.
  45. 45.J. Xu, S. Ramos, D. Vazquez, and A. M. Lopez. Hierarchical adaptive structural SVM for domain adaptation. 2016.

Citation

MLA
Ros, G., et al. “The SYNTHIA Dataset: A Large Collection of Synthetic Images for Semantic Segmentation of Urban Scenes”. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 3234–43, https://doi.org/10.1109/CVPR.2016.352.
APA
Ros, G., Sellart, L., Materzynska, J., Vazquez, D., & Lopez, A. M. (2016). The SYNTHIA Dataset: A Large Collection of Synthetic Images for Semantic Segmentation of Urban Scenes. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 3234–3243. https://doi.org/10.1109/CVPR.2016.352
Chicago
Ros, G., L. Sellart, J. Materzynska, D. Vazquez, and A. M. Lopez. 2016. “The SYNTHIA Dataset: A Large Collection of Synthetic Images for Semantic Segmentation of Urban Scenes”. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 3234–43. https://doi.org/10.1109/CVPR.2016.352.
Harvard
Ros, G. et al. (2016) “The SYNTHIA Dataset: A Large Collection of Synthetic Images for Semantic Segmentation of Urban Scenes”, 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 3234–3243. Available at: https://doi.org/10.1109/CVPR.2016.352.
Vancouver
1. Ros G, Sellart L, Materzynska J, Vazquez D, Lopez AM (2016) The SYNTHIA Dataset: A Large Collection of Synthetic Images for Semantic Segmentation of Urban Scenes. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 3234–3243

BibTeX

@inproceedings{Ros_2016, title={The SYNTHIA Dataset: A Large Collection of Synthetic Images for Semantic Segmentation of Urban Scenes}, url={http://dx.doi.org/10.1109/CVPR.2016.352}, DOI={10.1109/cvpr.2016.352}, booktitle={2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Ros, German and Sellart, Laura and Materzynska, Joanna and Vazquez, David and Lopez, Antonio M.}, year={2016}, month=June, pages={3234–3243} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: IEEE