3D-VField: Adversarial Augmentation of Point Clouds for Domain Generalization in 3D Object Detection

Alexander LehnerStefano GasperiniAlvaro Marcos-RamiroMichael SchmidtMohammad-Ali Nikouei MahaniNassir NavabBenjamin BusamFederico Tombari

article2022CVPR64 citations

Proposes 3D-VField, a sensor-aware data augmentation framework that deforms point clouds along sensor view rays using adversarially learned vector fields to improve 3D object detection generalization on out-of-domain and rare object geometries, supported by a new crash scenario dataset.

Listen

Three-dimensional object detection using spatial point cloud sensors, such as light detection and ranging (LiDAR), is critical for the safety of autonomous driving and robotics. Standard detection models rely heavily on learned geometric point relationships and frequently fail when encountering non-standard, rare, or damaged vehicles. Because rare real-world edge cases are sparsely represented in standard training sets, these natural variations act as adversarial samples that produce dangerous false negatives and false positives in deployment.

The article introduces and evaluates 3D-VField, a sensor-aware data augmentation framework designed to improve the generalization of 3D object detectors to unseen and out-of-domain environments. The method generates plausible, transferable geometric deformations during model training without adding or removing points, and it introduces CrashD, a public synthetic dataset featuring damaged and rare vehicles, to benchmark out-of-domain detection robustness.

The authors train sample-independent 3D vector fields using adversarial learning against detection networks, optimizing vectors that shift points while constraining movements strictly along the sensor's line of sight and smoothing shifts across adjacent surfaces. To evaluate real-world transferability, detectors trained on the standard German KITTI dataset were tested without fine-tuning on diverse out-of-domain targets: the real-world United States Waymo Open Dataset, the synthetic CrashD dataset containing 46,936 clean and damaged vehicles across 15,340 scenes, and indoor benchmark data.

The evaluation yielded several key findings. First, existing detectors experience severe performance degradation on non-standard objects; on the baseline PointPillars model, average precision dropped from 65.20% on normal clean vehicles to 22.48% on rare crashed vehicles. Second, training models with 3D-VField significantly improved out-of-domain detection accuracy, achieving a 9% relative gain over standard training on the Waymo dataset and raising detection of rare crashed vehicles on CrashD to 30.37% without sacrificing in-domain accuracy on KITTI. Third, combining 3D-VField with existing domain adaptation techniques tripled detection accuracy on the most challenging rare crash cases compared to the baseline detector. Finally, the approach proved architecture-agnostic, consistently enhancing performance across multiple distinct 3D detection architectures including PointPillars, Second, and Part-A2.

These findings demonstrate that perception failures on rare or deformed shapes stem from training distribution gaps that cannot be solved by standard geometric augmentations like flips or rotations alone. Plausible, sensor-constrained synthetic deformations provide a practical, low-cost training strategy to reduce high-stakes perception risks in autonomous systems. Furthermore, the results indicate that data augmentation and target domain adaptation are complementary, offering substantial safety enhancements when paired.

Engineering and safety teams should integrate sensor-aware adversarial point cloud augmentations into existing perception training pipelines to improve corner-case detection. Organizations developing autonomous systems should also incorporate structured out-of-domain benchmarks, such as the public CrashD dataset, into their safety verification workflows to systematically audit detector robustness against damaged and atypical vehicles before road deployment.

Confidence in these findings is high across standard autonomous driving benchmarks and distinct sensor modalities. However, limitations remain: the core evaluation transfers models trained in Germany (KITTI) to simulated crashes (CrashD) and US driving conditions (Waymo), meaning performance gains may vary in dynamic real-world crash scenes featuring unpredictable physical debris and severe structural occlusions.

Cover for 3D-VField: Adversarial Augmentation of Point Clouds for Domain Generalization in 3D Object Detection

Abstract

As 3D object detection on point clouds relies on the geometrical relationships between the points, non-standard object shapes can hinder a method's detection capability. However, in safety-critical settings, robustness to out-of-domain and long-tail samples is fundamental to circumvent dangerous issues, such as the misdetection of damaged or rare cars. In this work, we substantially improve the generalization of 3D object detectors to out-of-domain data by deforming point clouds during training. We achieve this with 3D-VField: a novel data augmentation method that plausibly deforms objects via vector fields learned in an adversarial fashion. Our approach constrains 3D points to slide along their sensor view rays while neither adding nor removing any of them. The obtained vectors are transferable, sample-independent and preserve shape and occlusions. Despite training only on a standard dataset, such as KITTI, augmenting with our vector fields significantly improves the generalization to differently shaped objects and scenes. Towards this end, we propose and share CrashD: a synthetic dataset of realistic damaged and rare cars, with a variety of crash scenarios. Extensive experiments on KITTI, Waymo, our CrashD and SUN RGB-D show the generalizability of our techniques to out-of-domain data, different models and sensors, namely LiDAR and ToF cameras, for both indoor and outdoor scenes. Our CrashD dataset is available at https://crashd-cars.github.io.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Improving Generalization
  • 2.1.1 Generalization for 3D Object Detection
  • 2.2. Adversarial Examples
  • 2.2.1 Adversarial point clouds
  • 3. Method
  • 3.1. Adversarially learned vector field
  • 3.2. Objects Deformation
  • 3.3. Adversarial Data Augmentation
  • 4. Experiments and Results
  • 4.1. Experimental Setup
  • 4.2. Quantitative Results
  • 4.3. Qualitative Results
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Sensor-aware transferable vector fields for adversarial augmentation

    model/method

    3D-VField is a domain-generalization method that learns a compact, sample-independent vector field and uses it to deform 3D object point clouds during detector training. A vector field is represented by displacement vectors attached to a uniformly spaced lattice inside a default 3D object bounding box. The same learned field can therefore be transferred to previously seen or unseen objects after scaling it to the target object's dimensions.

    The deformation preserves the number of points and does not add or remove geometry. Each point is displaced only along the viewing ray from the 3D sensor to that point, so the transformed cloud remains compatible with the sensor's acquisition geometry and preserves the object's occlusion structure. The resulting deformed objects are used as adversarial training augmentations intended to resemble plausible out-of-domain shape variations rather than maximally destructive attacks.

  2. Knowl 2 — Adversarial objective for learning the vector fields

    equation

    For each default object bounding box BoB_o, 3D-VField places a lattice of displacement vectors with spacing tt. The box is parameterized by width ww, height hh, length ll, orientation angle α\alpha, and center c=(x,y,z)c=(x,y,z). During field learning, the same candidate vector field is applied to target objects across the training scenes and optimized against a fixed 3D detector.

    A predicted bounding-box proposal qq with confidence score s∈(0,1)s\in(0,1) is considered relevant when s>0.1s>0.1. Let QQ be the set of relevant proposals, and let q∗q^\ast be the ground-truth box corresponding to a proposal. The field minimizes the confidence of relevant proposals, weighted by their 3D intersection-over-union with the ground truth:

    Ladv=∑(q,s)∈Q−IoU⁡(q∗,q)log⁡(1−s).\mathcal{L}_{\mathrm{adv}}=\sum_{(q,s)\in Q}-\operatorname{IoU}(q^\ast,q)\log(1-s).

    Minimizing this loss encourages the detector either to miss the object or to produce a misaligned box. After the loss converges, the learned vectors are frozen and reused as data augmentations.

  3. Knowl 3 — Ray-constrained and smooth point deformation

    model/method

    Before deforming an object, 3D-VField scales the learned lattice to the target object's size. For every object point pip_i, the method computes the optical ray uiu_i from the sensor to pip_i. The displacement vector attached to the nearest lattice location is projected onto uiu_i, producing a ray-consistent shift; the point is moved only by this projected component. Thus, no point changes its sensor-view direction, and no point is inserted or deleted.

    The displacement vectors are bounded by an L∞L_\infty constraint ∥v∥∞<ϵ\lVert v\rVert_\infty<\epsilon. To avoid abrupt changes along the object's surface, the method uses the kk nearest lattice vectors for each point. Let vijv_{ij} be the jjth nearest lattice vector for point pip_i, let rijr_{ij} be its sensor-ray projection, and let dijd_{ij} be the Euclidean distance between pip_i and the lattice location associated with vijv_{ij}. The final displacement mim_i is the distance-weighted average

    mi=∑j=1kdijrijk.m_i=\frac{\sum_{j=1}^{k}d_{ij}r_{ij}}{k}.

    Averaging neighboring projected vectors makes opposite local directions cancel and produces smoother depth changes across the object surface. The default experiments used k=2k=2 and ϵ=30 cm\epsilon=30\,\mathrm{cm}.

  4. Knowl 4 — Rotation-grouped field learning and detector-training augmentation

    algorithm

    Input: labeled 3D training scenes, a 3D detector, default box dimensions w=1.8 mw=1.8\,\mathrm{m}, h=1.6 mh=1.6\,\mathrm{m}, l=4.6 ml=4.6\,\mathrm{m}, lattice spacing t=20 cmt=20\,\mathrm{cm}, G=12G=12 relative-orientation groups, and N=6N=6 fields per group.

    Output: a detector trained with normal and 3D-VField-deformed objects.

    1. Partition training objects into GG groups according to the relative orientation between each object and the sensor.
    2. For every orientation group and every one of its NN fields, initialize the lattice displacement vectors independently from a uniform distribution in [−1 cm,1 cm][-1\,\mathrm{cm},1\,\mathrm{cm}].
    3. Apply a candidate field to every target object in the training scenes, scale it to the object size, project its vectors onto the sensor rays, smooth the point shifts using the k=2k=2 nearest vectors, and clamp the vector magnitudes using ϵ=30 cm\epsilon=30\,\mathrm{cm}.
    4. Backpropagate the adversarial detector loss through the deformed point clouds and update the field vectors with Adam at learning rate 0.050.05. Continue until the adversarial loss converges.
    5. Freeze all learned fields. During ordinary detector training, randomly select one object in each scene, randomly select one of the NN fields associated with that object's relative-orientation group, and apply the corresponding sensor-aware deformation.
    6. Train the detector on the resulting mixture of undeformed and deformed scenes. Randomly selecting one object and one field per scene exposes the detector to many structurally consistent variations without forcing it to memorize one fixed deformation.
  5. Knowl 5 — CrashD benchmark for rare and damaged vehicles

    data/table

    CrashD is a publicly released synthetic out-of-domain benchmark designed to test 3D object detectors on plausible vehicle-shape changes. It contains normal, old, sports, and damaged cars generated with a realistic vehicle simulator. Damage is characterized by intensity—light, moderate, or hard—and by type: clean undamaged vehicles, linear frontal or rear damage, and lateral T-bone damage.

    The dataset contains 15,340 automatically generated scenes captured with a 64-beam LiDAR configured to mimic KITTI. Each scene contains between 1 and 5 vehicles, and the collection contains 46,936 cars in total. Damaged vehicles were captured before repair and then repaired and placed at the same locations to produce corresponding clean examples. This paired construction isolates the effect of damage while keeping scene placement fixed; rare vehicle types and damaged vehicles provide increasingly difficult shifts away from the normal cars used for training.

  6. Knowl 6 — Evaluation protocol for cross-domain 3D detection

    experimental setup

    The main source-domain experiments use the standard KITTI split containing 3,712 training and 3,769 validation LiDAR point clouds, with the car class evaluated at the easy, moderate, and hard levels. Detectors trained on KITTI are transferred to Waymo and CrashD without fine-tuning. The method is also applied to SUN RGB-D to test applicability to indoor time-of-flight point clouds.

    The evaluated detectors are PointPillars, SECOND, and Part-A2 for outdoor detection, together with VoteNet for indoor detection. Detection quality is measured by average precision using a 3D IoU threshold of 0.70.7 on KITTI and CrashD, 0.50.5 on Waymo, and 0.250.25 on SUN RGB-D. Attack success rate (ASR) is the percentage of objects that become false negatives after deformation, where an object is counted as detected when its 3D IoU exceeds 0.70.7.

    For comparisons with other adversarial augmentations, all methods use the same KITTI split, a maximum perturbation of 30 cm30\,\mathrm{cm}, and random augmentation of one object per scene. The compared alternatives include iterative-gradient L2L_2 perturbation, Chamfer perturbation, adversarial point generation that adds 10%10\% of the points, and adversarial point removal that removes 10%10\% of the points. The experiments also combine 3D-VField with statistical-normalization domain adaptation, which rescales KITTI objects using average target-domain box dimensions and fine-tunes on the altered source data.

  7. Knowl 7 — Cross-domain detection results on KITTI, Waymo, and CrashD

    data/table

    The comparison below reports KITTI AP, attack success rate, and transfer AP from KITTI to Waymo and CrashD without fine-tuning. The PointPillars rows compare standard augmentation and adversarial alternatives under the same training protocol; the SECOND and Part-A2 rows test whether 3D-VField transfers across detector architectures. For CrashD, normal and rare refer to vehicle type, while clean and crash refer to the undamaged and damaged versions. Asterisks mark sample-specific attacks whose ASR was learned on the KITTI validation set rather than transferred from a generic field.

    Could not parse LaTeX table

    For PointPillars, 3D-VField raises Waymo AP from 40.8640.86 to 44.6144.61 without reducing KITTI moderate AP relative to the baseline. It also gives the best transfer AP among the adversarial augmentations across the normal and rare CrashD categories. Rare crash vehicles are the hardest condition: PointPillars reaches 22.4822.48 AP, whereas 3D-VField reaches 30.3730.37. Combining 3D-VField with target-aware statistical normalization further raises Waymo AP to 51.3251.32 and rare-crash AP to 76.4276.42.

  8. Knowl 8 — Transferability across detector architectures and domain-adaptation strategies

    empirical result

    The vector fields were learned against PointPillars, but their augmentations improved the out-of-domain performance of all three evaluated outdoor detectors. For SECOND, Waymo AP increased from 42.4542.45 to 43.5143.51, normal-crash CrashD AP from 56.7456.74 to 60.5160.51, and rare-crash AP from 32.8432.84 to 36.1436.14. For Part-A2, Waymo AP increased from 49.7649.76 to 56.0856.08, normal-crash AP from 63.2563.25 to 73.8073.80, and rare-crash AP from 52.3352.33 to 61.3461.34. KITTI AP changed by at most a small amount in these comparisons.

    3D-VField is a domain-generalization method and does not use target-domain data. It is therefore complementary to statistical-normalization domain adaptation. When the target-aware size normalization is combined with 3D-VField, Waymo AP rises from 49.2749.27 with statistical normalization alone to 51.3251.32. On CrashD, the combined method reaches 87.2887.28 AP for normal crash vehicles and 76.4276.42 for rare crash vehicles, compared with 72.5972.59 and 48.2348.23 for statistical normalization alone. The combined augmentation and adaptation strategy therefore uses the deformation diversity of 3D-VField in addition to target-domain size information.

  9. Knowl 9 — Ablations of deformation constraints and orientation grouping

    data/table

    All ablations below train on KITTI with a 30 cm30\,\mathrm{cm} deformation limit and evaluate transfer without fine-tuning. The constraint ablation isolates the effect of learning, sensor-ray consistency, and surface smoothness. The orientation-grouping ablation measures the trade-off between attack specificity, transfer generalization, and the number of stored vectors.

    Could not parse LaTeX table

    The unconstrained learned field has a very high ASR but transfers poorly to Waymo and normal CrashD vehicles. Restricting motion to sensor rays improves the difficult rare-crash transfer, and adding distance-based smoothing gives the best overall transfer while preserving KITTI performance.

    Could not parse LaTeX table

    Using G=12G=12 relative-orientation groups gives the strongest Waymo transfer and a good balance between generality and attack specificity. Increasing the grouping to G=360G=360 greatly increases storage and overfits the training objects, reducing validation ASR and Waymo AP. With G=12G=12 and N=6N=6, the method stores about 120,000120\\,000 vectors, compared with 10.910.9 million and 12.612.6 million vectors required by the sample-specific iterative-gradient and Chamfer attacks for their training and validation sets.

  10. Knowl 10 — Plausibility–attack-strength trade-off and observed failure modes

    limitation

    3D-VField deliberately does not maximize attack success rate. Its sample-independent fields achieve lower ASR than sample-specific attacks because they must remain transferable across objects and scenes. The paper argues that extremely high ASR often corresponds to unrecognizable point clouds and is not desirable for data augmentation; the useful regime alters object geometry enough to add training diversity while remaining close to plausible sensor observations.

    The deformation limit also creates a trade-off. Increasing the maximum displacement from 30 cm30\,\mathrm{cm} to 40 cm40\,\mathrm{cm} or 60 cm60\,\mathrm{cm} raises ASR to 73.3%73.3\% and 87.1%87.1\%, respectively, but decreases KITTI AP by 1%1\% and 1.7%1.7\%. The authors therefore use 30 cm30\,\mathrm{cm} as a compromise between deformation strength and plausibility.

    Qualitative comparisons show that Chamfer-based attacks can produce highly distorted, effectively unrecognizable cars, whereas 3D-VField produces smaller shape-preserving changes. On challenging damaged CrashD vehicles, 3D-VField yields more aligned detections than the iterative-gradient augmentation in the reported examples. On Waymo, all methods struggle with cars represented by very few points, so sensor-aware deformation does not eliminate failures caused by severe sparsity or large domain shifts.

Coverage note — Supplementary-only numerical results for SUN RGB-D, noise robustness, detailed CrashD breakdowns, and additional grouping or aggregation ablations were omitted because their values are not present in the supplied main-paper text.

References

  1. 1.Rima Alaifari, Giovanni S. Alberti, and Tandri Gauksson. ADef: An iterative algorithm to construct adversarial deformations. In Proceedings of the International Conference on Learning Representations, 2019. 2, 3
  2. 2.Isabela Albuquerque, Nikhil Naik, Junnan Li, Nitish Keskar, and Richard Socher. Improving out-of-distribution generalization via multi-task self-supervised pretraining. arXiv preprint arXiv:2003.13525, 2020. 2
  3. 3.Yogesh Balaji, Swami Sankaranarayanan, and Rama Chellappa. Metareg: Towards domain generalization using meta-regularization. Advances in Neural Information Processing Systems, 31:998–1008, 2018. 2
  4. 4.Sara Beery, Yang Liu, Dan Morris, Jim Piavis, Ashish Kapoor, Neel Joshi, Markus Meister, and Pietro Perona. Synthetic examples improve generalization for rare classes. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 863–873, 2020. 1, 2
  5. 5.Daniel Bogdoll, Jasmin Breitenstein, Florian Heidecker, Maarten Bieshaar, Bernhard Sick, Tim Fingscheidt, and Marius Zollner. Description of corner cases in automated driving: Goals and challenges. In IEEE/CVF International Conference on Computer Vision Workshop, pages 1023–1028, 2021. 1
  6. 6.Yulong Cao, Ningfei Wang, Chaowei Xiao, Dawei Yang, Jin Fang, Ruigang Yang, Qi Alfred Chen, Mingyan Liu, and Bo Li. Invisible for both camera and LiDAR: Security of multi-sensor fusion based perception in autonomous driving under physical-world attacks. In Proceedings of the IEEE Symposium on Security and Privacy, pages 176–194, 2021. 3
  7. 7.Yulong Cao, Chaowei Xiao, Benjamin Cyr, Yimeng Zhou, Won Park, Sara Rampazzi, Qi Alfred Chen, Kevin Fu, and Z Morley Mao. Adversarial sensor attack on LiDAR-based perception in autonomous driving. In Proceedings of the 2019 ACM SIGSAC Conference on Computer and Communications Security, pages 2267–2281, 2019. 3
  8. 8.Yulong Cao, Chaowei Xiao, Dawei Yang, Jing Fang, Ruigang Yang, Mingyan Liu, and Bo Li. Adversarial objects against lidar-based autonomous driving systems. arXiv preprint arXiv:1907.05418, 2019. 3
  9. 9.Nicholas Carlini and David Wagner. Towards evaluating the robustness of neural networks. In Proceedings of the IEEE Symposium on Security and Privacy, pages 39–57, 2017. 2
  10. 10.MMDetection3D Contributors. MMDetection3D: OpenMMLab next-generation platform for general 3D object detection. https://github.com/open-mmlab/mmdetection3d, 2020. 5, 7
  11. 11.Stefano Gasperini, Jan Haug, Mohammad-Ali Nikouei Mahani, Alvaro Marcos-Ramiro, Nassir Navab, Benjamin Busam, and Federico Tombari. CertainNet: Sampling-free uncertainty estimation for object detection. IEEE Robotics and Automation Letters, 7(2):698–705, 2021. 1, 2
  12. 12.Stefano Gasperini, Patrick Koch, Vinzenz Dallabetta, Nassir Navab, Benjamin Busam, and Federico Tombari. R4Dyn: Exploring radar for self-supervised monocular depth estimation of dynamic scenes. In Proceedings of the IEEE International Conference on 3D Vision (3DV), pages 751–760, 2021. 1
  13. 13.Andreas Geiger, Philip Lenz, and Raquel Urtasun. Are we ready for autonomous driving? the KITTI vision benchmark suite. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3354–3361. IEEE, 2012. 1, 2, 5, 6, 8
  14. 14.Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and harnessing adversarial examples. In Proceedings of the International Conference on Learning Representations, 2015. 2
  15. 15.Abdullah Hamdi, Sara Rojas, Ali Thabet, and Bernard Ghanem. AdvPC: Transferable adversarial perturbations on 3D point clouds. In Proceedings of the European Conference on Computer Vision, pages 241–257. Springer, 2020. 3
  16. 16.Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, et al. The many faces of robustness: A critical analysis of out-of-distribution generalization. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8340–8349, 2021. 1, 2
  17. 17.Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt, and Dawn Song. Natural adversarial examples. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15262–15271, 2021. 1, 2
  18. 18.Alex H Lang, Sourabh Vora, Holger Caesar, Lubing Zhou, Jiong Yang, and Oscar Beijbom. PointPillars: Fast encoders for object detection from point clouds. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12697–12705, 2019. 1, 5, 6, 7, 8
  19. 19.Daniel Liu, Ronald Yu, and Hao Su. Adversarial shape perturbations on 3D point clouds. In Proceedings of the European Conference on Computer Vision, pages 88–104. Springer, 2020. 2, 3, 5, 6, 7, 8
  20. 20.Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. Towards deep learning models resistant to adversarial attacks. In Proceedings of the International Conference on Learning Representations, 2018. 3, 4
  21. 21.Pascale Maul, Marc Mueller, Fabian Enkler, Eva Pigova, Thomas Fischer, and Lefteris Stamatogiannakis. BeamNG.tech technical paper, 2021. 5
  22. 22.Jisoo Mok, Byunggook Na, Hyeokjun Choe, and Sungroh Yoon. AdvRush: Searching for adversarially robust neural architectures. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 12322–12332, 2021. 1, 2
  23. 23.Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, and Pascal Frossard. DeepFool: A simple and accurate method to fool deep neural networks. In IEEE Conference on Computer Vision and Pattern Recognition, pages 2574–2582. IEEE, 2016. 2
  24. 24.Nicolas Papernot, Patrick McDaniel, Somesh Jha, Matt Fredrikson, Z Berkay Celik, and Ananthram Swami. The limitations of deep learning in adversarial settings. In Proceedings of the IEEE European Symposium on Security and Privacy, pages 372–387, 2016. 2
  25. 25.Charles R Qi, Or Litany, Kaiming He, and Leonidas J Guibas. Deep Hough voting for 3D object detection in point clouds. In Proceedings of the IEEE International Conference on Computer Vision, 2019. 5
  26. 26.Fengchun Qiao, Long Zhao, and Xi Peng. Learning to learn single domain generalization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12556–12565, 2020. 1, 2
  27. 27.Martin Rabe, Stefan Milz, and Patrick Mader. Development methodologies for safety critical machine learning applications in the automotive domain: A survey. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 129–141, 2021. 1
  28. 28.Shaoshuai Shi, Zhe Wang, Jianping Shi, Xiaogang Wang, and Hongsheng Li. From points to parts: 3D object detection from point cloud with part-aware and part-aggregation network. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020. 5, 6, 7
  29. 29.Andrea Simonelli, Samuel Rota Bulo, Lorenzo Porzi, Elisa Ricci, and Peter Kontschieder. Towards generalization across depth for monocular 3D object detection. In Proceedings of the European Conference on Computer Vision, pages 767–782. Springer, 2020. 2
  30. 30.Shuran Song, Samuel P Lichtenberg, and Jianxiong Xiao. SUN RGB-D: A RGB-D scene understanding benchmark suite. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 567–576, 2015. 2, 5
  31. 31.Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: A simple way to prevent neural networks from overfitting. The Journal of Machine Learning Research, 15(1):1929–1958, 2014. 2
  32. 32.Cecilia Summers and Michael J Dinneen. Improved mixed-example data augmentation. In Proceedings of the IEEE Winter Conference on Applications of Computer Vision, pages 1262–1270, 2019. 2
  33. 33.Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, et al. Scalability in perception for autonomous driving: Waymo open dataset. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2446–2454, 2020. 2, 5, 6, 8
  34. 34.Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. Intriguing properties of neural networks. In Proceedings of the International Conference on Learning Representations, 2014. 2
  35. 35.James Tu, Mengye Ren, Sivabalan Manivasagam, Ming Liang, Bin Yang, Richard Du, Frank Cheng, and Raquel Urtasun. Physically realizable adversarial examples for LiDAR object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13713–13722, 2020. 1, 2, 3, 4, 5
  36. 36.Riccardo Volpi, Hongseok Namkoong, Ozan Sener, John Duchi, Vittorio Murino, and Silvio Savarese. Generalizing to unseen domains via adversarial data augmentation. In Proceedings of the International Conference on Neural Information Processing Systems, pages 5339–5349, 2018. 2
  37. 37.Jindong Wang, Cuiling Lan, Chang Liu, Yidong Ouyang, and Tao Qin. Generalizing to unseen domains: A survey on domain generalization. In Proceedings of the International Joint Conference on Artificial Intelligence, pages 4627–4635, 2021. 1, 2
  38. 38.Run Wang, Felix Juefei-Xu, Qing Guo, Yihao Huang, Xiaofei Xie, Lei Ma, and Yang Liu. Amora: Black-box adversarial morphing attack. In Proceedings of the ACM International Conference on Multimedia, pages 1376–1385, 2020. 3
  39. 39.Yan Wang, Xiangyu Chen, Yurong You, Li Erran Li, Bharath Hariharan, Mark Campbell, Kilian Q Weinberger, and Wei-Lun Chao. Train in Germany, test in the USA: Making 3D object detectors generalize. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11713–11723, 2020. 1, 2, 4, 5, 6, 7
  40. 40.Chong Xiang, Charles R. Qi, and Bo Li. Generating 3D adversarial point clouds. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9128–9136, 2019. 2, 3, 5, 6, 7, 8
  41. 41.Chaowei Xiao, Bo Li, Jun-yan Zhu, Warren He, Mingyan Liu, and Dawn Song. Generating Adversarial Examples with Adversarial Networks. In Proceedings of the International Joint Conference on Artificial Intelligence, pages 3905–3911, July 2018. 2
  42. 42.Yan Yan, Yuxing Mao, and Bo Li. SECOND: Sparsely Embedded Convolutional Detection. Sensors, 18(10):3337, Oct. 2018. 5, 6, 7
  43. 43.Jiancheng Yang, Qiang Zhang, Rongyao Fang, Bingbing Ni, Jinxian Liu, and Qi Tian. Adversarial attack and defense on point sets. arXiv preprint arXiv:1902.10899, 2019. 3, 5, 6, 7
  44. 44.Xiaoyong Yuan, Pan He, Qile Zhu, and Xiaolin Li. Adversarial Examples: Attacks and Defenses for Deep Learning. IEEE Transactions on Neural Networks and Learning Systems, 30(9):2805–2824, 2019. 2
  45. 45.Linjun Zhang, Zhun Deng, Kenji Kawaguchi, Amirata Ghorbani, and James Zou. How does mixup help with robustness and generalization? In Proceedings of the International Conference on Learning Representations, 2021. 2

Citation

MLA
Lehner, A., et al. “3D-VField: Adversarial Augmentation of Point Clouds for Domain Generalization in 3D Object Detection”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 17274–83, https://doi.org/10.1109/CVPR52688.2022.01678.
APA
Lehner, A., Gasperini, S., Marcos-Ramiro, A., Schmidt, M., Mahani, M.-A. N., Navab, N., Busam, B., & Tombari, F. (2022). 3D-VField: Adversarial Augmentation of Point Clouds for Domain Generalization in 3D Object Detection. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 17274–17283. https://doi.org/10.1109/CVPR52688.2022.01678
Chicago
Lehner, A., S. Gasperini, A. Marcos-Ramiro, et al. 2022. “3D-VField: Adversarial Augmentation of Point Clouds for Domain Generalization in 3D Object Detection”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 17274–83. https://doi.org/10.1109/CVPR52688.2022.01678.
Harvard
Lehner, A. et al. (2022) “3D-VField: Adversarial Augmentation of Point Clouds for Domain Generalization in 3D Object Detection”, 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 17274–17283. Available at: https://doi.org/10.1109/CVPR52688.2022.01678.
Vancouver
1. Lehner A, Gasperini S, Marcos-Ramiro A, Schmidt M, Mahani M-AN, Navab N, Busam B, Tombari F (2022) 3D-VField: Adversarial Augmentation of Point Clouds for Domain Generalization in 3D Object Detection. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 17274–17283

BibTeX

@inproceedings{Lehner_2022, title={3D-VField: Adversarial Augmentation of Point Clouds for Domain Generalization in 3D Object Detection}, url={http://dx.doi.org/10.1109/CVPR52688.2022.01678}, DOI={10.1109/cvpr52688.2022.01678}, booktitle={2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Lehner, Alexander and Gasperini, Stefano and Marcos-Ramiro, Alvaro and Schmidt, Michael and Mahani, Mohammad-Ali Nikouei and Navab, Nassir and Busam, Benjamin and Tombari, Federico}, year={2022}, month=June, pages={17274–17283} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE