Fast Point Cloud Generation with Straight Flows

Lemeng WuDilin WangChengyue GongXingchao LiuYunyang XiongRakesh RanjanRaghuraman KrishnamoorthiVikas ChandraQiang Liu

article2023CVPR74 citations

Proposes Point Straight Flow, a novel framework that straightens generative transport trajectories and distills them into a single step, enabling high-quality 3D point cloud generation over 700 times faster than standard diffusion models.

Listen

Generating high-quality 3D point clouds is critical for real-world technologies such as autonomous driving, robotics, and virtual reality. While diffusion models produce state-of-the-art, realistic 3D shapes, they require simulating thousands of iterative denoising steps. This iterative process creates severe computational bottlenecks and high latency, making diffusion models impractical for time-sensitive, production-level deployment.

The article demonstrates and evaluates Point Straight Flow (PSF), a generative framework designed to produce high-quality 3D point clouds in a single step. The primary objective is to drastically reduce generation latency while matching the shape quality of standard, multi-step diffusion models.

The authors develop a three-stage training framework. First, they train an initial velocity flow model using continuous differential equations rather than random noise-driven processes. Second, they straighten the transport trajectory using a reflow optimization technique that minimizes transport costs. Third, they distill this straightened trajectory into a one-step generator, utilizing Chamfer distance—a metric tailored to irregular, unordered 3D point sets—to preserve geometric structure. The framework was evaluated on standard benchmark datasets across unconditional 3D shape generation, 3D point cloud completion, and text-guided shape synthesis on modern graphics hardware.

The evaluation yielded several key findings. First, PSF generated realistic 3D point clouds in approximately 0.04 seconds per sample, achieving over a 700-fold speedup compared to standard 1,000-step diffusion baselines and over a 75-fold speedup compared to 100-step accelerated diffusion models. Second, this extreme speedup was attained with negligible loss in sample quality and geometric fidelity across standard categories including airplanes, chairs, and cars. Third, in training-free, text-guided shape generation, PSF completed the synthesis process in 12 seconds compared to roughly 15 minutes for standard diffusion methods. Fourth, when applied to autonomous vehicle sensor pipelines, PSF completed sparse LiDAR point clouds across multiple vehicles in 0.2 seconds, well within real-time operating constraints.

These findings indicate that generative 3D modeling can transition from slow, offline rendering to low-latency, real-time edge environments. By demonstrating that straight transport paths can be compressed into a single neural evaluation without degrading structural fidelity, PSF removes computational cost and runtime barriers in automated 3D perception and simulation workflows.

Engineering and research teams developing real-time 3D perception systems, such as autonomous driving perception stacks or interactive design tools, should evaluate straight-flow formulations as replacements for standard multi-step diffusion pipelines. Further validation should focus on deploying PSF across larger-scale outdoor scenes, highly dense point clouds, and diverse embedded hardware environments to verify stability under variable compute constraints.

Cover for Fast Point Cloud Generation with Straight Flows

Abstract

Diffusion models have emerged as a powerful tool for point cloud generation. A key component that drives the impressive performance for generating high-quality samples from noise is iteratively denoise for thousands of steps. While beneficial, the complexity of learning steps has limited its applications to many 3D real-world. To address this limitation, we propose Point Straight Flow (PSF), a model that exhibits impressive performance using one step. Our idea is based on the reformulation of the standard diffusion model, which optimizes the curvy learning trajectory into a straight path. Further, we develop a distillation strategy to shorten the straight path into one step without a performance loss, enabling applications to 3D real-world with latency constraints. We perform evaluations on multiple 3D tasks and find that our PSF performs comparably to the standard diffusion model, outperforming other efficient 3D point cloud generation methods. On real-world applications such as point cloud completion and training-free text-guided generation in a low-latency setup, PSF performs favorably.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 2.1. Generative model with transport flow
  • 2.2. Fast sampling for transport model
  • 2.3. 3D generative model
  • 3. Point Straight Flow
  • 4. Experiment
  • 4.1. Unconditional point cloud generation
  • 4.2. Training-free text-guided shape generation
  • 4.3. Point Cloud Completion
  • 4.4. Ablation Study
  • 5. Discussion and Conclusion
  • References

Knowls

  1. Knowl 1 — Point Straight Flow Framework

    model/method

    Point Straight Flow (PSF) is a generative framework designed to produce 3D point clouds in a single step by straightening the transport trajectory between a standard Gaussian prior and the target shape distribution.

    Let a 3D point cloud be denoted as X∈RM×3X \in \mathbb{R}^{M \times 3}, where MM is the number of points. Let X0∼N(0,I)X_0 \sim \mathcal{N}(0, I) be Gaussian noise and X1∼DX_1 \sim \mathcal{D} be a real point cloud from dataset D\mathcal{D}. The generative transport is formulated as an ordinary differential equation (ODE):

    dXtdt=vθ(Xt,t),t∈[0,1]\frac{\mathrm{d}X_t}{\mathrm{d}t} = v_\theta(X_t, t), \quad t \in [0, 1]

    where Xt=tX1+(1−t)X0X_t = t X_1 + (1 - t) X_0 represents the linear interpolation at time t∈[0,1]t \in [0, 1], and vθ:RM×3×[0,1]→RM×3v_\theta: \mathbb{R}^{M \times 3} \times [0, 1] \to \mathbb{R}^{M \times 3} is a parameterized neural velocity field network.

    PSF consists of a three-stage training process:

    1. Initial Velocity Flow Training: The network vθv_\theta is trained to approximate the straight-line displacement direction X1−X0X_1 - X_0 via regression across random time steps t∼U(0,1)t \sim \mathcal{U}(0, 1).

    2. Reflow for Trajectory Straightening: Fixed point cloud pairs (X0′,X1′)(X'_0, X'_1) are generated by integrating the trained ODE model using an Euler solver with N=1000N = 1000 steps starting from noise X0′∼N(0,I)X'_0 \sim \mathcal{N}(0, I). The network vθv_\theta is fine-tuned on these paired endpoints (X0′,X1′)(X'_0, X'_1), which reduces the transport cost and enforces straight velocity trajectories.

    3. One-Step Flow Distillation: The straightened trajectory model is distilled into a single-step generator X1′=X0′+vθ(X0′,0)X'_1 = X'_0 + v_\theta(X'_0, 0) by minimizing the Chamfer Distance between the one-step prediction and the NN-step integration target X1′X'_1.

  2. Knowl 2 — Point Straight Flow Training and Sampling Algorithm

    algorithm

    The Point Straight Flow (PSF) pipeline trains an initial velocity field, straightens its transport trajectories via reflow, and distills the model into a one-step point cloud generator.

    The neural network architecture vθv_\theta uses a Point-Voxel CNN (PVCNN) styled U-Net. In Stage 1, the model is trained for 200k steps with a batch size of 256, learning rate 2×10−42 \times 10^{-4}, and exponential moving average (EMA) rate of 0.9999. In Stage 2 (Reflow), 50k point pairs (X0′,X1′)(X'_0, X'_1) are sampled using an Euler solver with N=1000N = 1000 steps, and vθv_\theta is fine-tuned for 10k steps with a learning rate of 2×10−52 \times 10^{-5}. In Stage 3 (Distillation), vθv_\theta is fine-tuned for another 10k steps with a learning rate of 2×10−52 \times 10^{-5} using Chamfer Distance.

    Input: Point cloud training dataset D\mathcal{D}, neural velocity field network vθv_\theta with parameters θ\theta, step count N=1000N = 1000
    Output: Trained one-step point cloud generator vθv_\theta
    // Stage 1: Training initial velocity flow model
    repeat
        Sample data point cloud X1∼DX_1 \sim \mathcal{D}, noise X0∼N(0,I)X_0 \sim \mathcal{N}(0, I), and t∼U(0,1)t \sim \mathcal{U}(0, 1)
        Construct Xt=tX1+(1−t)X0X_t = t X_1 + (1 - t) X_0
        Update θ\theta by minimizing ∥vθ(Xt,t)−(X1−X0)∥2\|v_\theta(X_t, t) - (X_1 - X_0)\|^2
    until Stage 1 convergence
    // Stage 2: Improving straightness via reflow
    Generate dataset S=∅\mathcal{S} = \emptyset
    for k=1k = 1 to 5000050000 do
        Sample X0′∼N(0,I)X'_0 \sim \mathcal{N}(0, I)
        Initialize Y0=X0′Y_0 = X'_0
        for t^=0\hat{t} = 0 to N−1N - 1 do
            Yt^+1=Yt^+1Nvθ(Yt^,t^N)Y_{\hat{t} + 1} = Y_{\hat{t}} + \frac{1}{N} v_\theta\left(Y_{\hat{t}}, \frac{\hat{t}}{N}\right)
        Set X1′=YNX'_1 = Y_N
        Add pair (X0′,X1′)(X'_0, X'_1) to S\mathcal{S}
    repeat
        Sample pair (X0′,X1′)∼S(X'_0, X'_1) \sim \mathcal{S} and t∼U(0,1)t \sim \mathcal{U}(0, 1)
        Construct Xt′=tX1′+(1−t)X0′X'_t = t X'_1 + (1 - t) X'_0
        Update θ\theta by minimizing ∥vθ(Xt′,t)−(X1′−X0′)∥2\|v_\theta(X'_t, t) - (X'_1 - X'_0)\|^2
    until Stage 2 convergence (10k iterations)
    // Stage 3: Flow distillation into one step
    repeat
        Sample pair (X0′,X1′)∼S(X'_0, X'_1) \sim \mathcal{S}
        Predict one-step shape X^1=X0′+vθ(X0′,0)\hat{X}_1 = X'_0 + v_\theta(X'_0, 0)
        Update θ\theta by minimizing Chamfer Distance CD(X^1,X1′)\text{CD}(\hat{X}_1, X'_1)
    until Stage 3 convergence (10k iterations)
    // Inference / Sampling
    Function SamplePointCloud():
        Sample X0∼N(0,I)X_0 \sim \mathcal{N}(0, I)
        return X0+vθ(X0,0)X_0 + v_\theta(X_0, 0)
  3. Knowl 3 — Permutation-Invariant Flow Distillation via Chamfer Distance

    equation

    In Point Straight Flow, one-step distillation compresses the ODE trajectory into a single update step X1′=X0′+vθ(X0′,0)X'_1 = X'_0 + v_\theta(X'_0, 0). Because 3D point clouds are unordered sets of coordinates, point-to-point Euclidean ℓ2\ell_2 regression fails to accommodate arbitrary point permutations between initial noise and target points. PSF defines the distillation objective using the symmetric Chamfer Distance:

    min⁡θE(X0′,X1′)[CD(X0′+vθ(X0′,0), X1′)]\min_\theta \mathbb{E}_{(X'_0, X'_1)} \left[ \text{CD}\left(X'_0 + v_\theta(X'_0, 0), \, X'_1\right) \right]

    where X0′∼N(0,I)X'_0 \sim \mathcal{N}(0, I) is the noise input, X1′X'_1 is the target point cloud generated by multi-step ODE simulation (N=1000N=1000), and the Chamfer Distance CD(Xi,Xj)\text{CD}(X_i, X_j) between two point clouds Xi,Xj⊂R3X_i, X_j \subset \mathbb{R}^3 is defined as:

    CD(Xi,Xj)=∑p∈Ximin⁡p^∈Xj∥p−p^∥2+∑p^∈Xjmin⁡p∈Xi∥p−p^∥2\text{CD}(X_i, X_j) = \sum_{p \in X_i} \min_{\hat{p} \in X_j} \|p - \hat{p}\|_2 + \sum_{\hat{p} \in X_j} \min_{p \in X_i} \|p - \hat{p}\|_2

  4. Knowl 4 — Reflow Formulation and Trajectory Straightness Metric

    equation

    The reflow procedure fine-tunes the velocity field network vθv_\theta on deterministic pairs (X0′,X1′)(X'_0, X'_1), where X0′∼N(0,I)X'_0 \sim \mathcal{N}(0, I) and X1′X'_1 is obtained by integrating the initial ODE using Euler discretization with N=1000N=1000 steps:

    X(t^+1)/N′=Xt^/N′+1Nvθ(Xt^/N′, t^N),t^∈{0,1,…,N−1}X'_{(\hat{t}+1)/N} = X'_{\hat{t}/N} + \frac{1}{N} v_\theta\left(X'_{\hat{t}/N}, \, \frac{\hat{t}}{N}\right), \quad \hat{t} \in \{0, 1, \dots, N-1\}

    Reflow minimizes the velocity regression loss over the interpolated states Xt′=tX1′+(1−t)X0′X'_t = t X'_1 + (1-t)X'_0 with t∼U(0,1)t \sim \mathcal{U}(0, 1):

    min⁡θEt∼U(0,1),(X0′,X1′)[∥vθ(Xt′,t)−(X1′−X0′)∥2]\min_\theta \mathbb{E}_{t \sim \mathcal{U}(0,1), (X'_0, X'_1)} \left[ \|v_\theta(X'_t, t) - (X'_1 - X'_0)\|^2 \right]

    The deviation of the simulated path from a perfectly straight line is quantified by the Straightness metric across NN discrete steps:

    Straightness=1N∑t^=0N−1∥(X1′−X0′)−vθ(Xt^/N′, t^N)∥2\text{Straightness} = \frac{1}{N} \sum_{\hat{t}=0}^{N-1} \left\| (X'_1 - X'_0) - v_\theta\left(X'_{\hat{t}/N}, \, \frac{\hat{t}}{N}\right) \right\|^2

    A straight path exhibits a constant velocity vector vθ(Xt^/N′,t^/N)=X1′−X0′v_\theta(X'_{\hat{t}/N}, \hat{t}/N) = X'_1 - X'_0 throughout the integration, achieving a Straightness score of 0.

  5. Knowl 5 — Unconditional 3D Point Cloud Generation Benchmark Results

    data/table

    Point Straight Flow (PSF) was evaluated on unconditional 3D shape generation using the ShapeNet dataset (split following PointFlow and PVD) across Airplane, Chair, and Car categories. Performance is assessed using 1-Nearest Neighbor Accuracy (1-NNA, where lower is better, optimal score is 50%) with Chamfer Distance (CD) and Earth Mover's Distance (EMD). Inference times were measured on an Nvidia RTX 3090 GPU with batch size 1 averaged over 50 trials.

    Model Sampling Time (s) Airplane Chair Car
    CD ↓\downarrow EMD ↓\downarrow CD ↓\downarrow EMD ↓\downarrow CD ↓\downarrow EMD ↓\downarrow
    1-GAN 0.03 87.30 93.95 68.58 83.84 66.49 88.78
    PointFlow 0.27 75.68 70.74 62.84 60.57 58.10 56.25
    DPF-Net 0.33 75.18 65.55 62.00 58.53 62.35 54.48
    SoftFlow 0.12 76.05 65.80 59.21 60.05 64.77 60.09
    SetVAE 0.03 75.31 77.65 58.76 61.48 59.66 61.48
    ShapeGF 0.34 80.00 76.17 68.96 65.48 63.20 56.53
    DPM 22.8 76.42 86.91 60.05 74.77 68.89 79.97
    PVD (N=1000N=1000) 29.9 73.82 64.81 56.26 53.32 54.55 53.83
    PVD-DDIM (N=100N=100) 3.15 76.21 69.84 61.54 57.73 60.95 59.35
    PSF (ours) 0.04 71.11 61.09 58.92 54.45 57.19 56.07

    PSF achieves sample quality competitive with the 1000-step diffusion baseline PVD while operating in a single step (0.04s per shape), yielding a speedup of over 700×700\times compared to PVD and over 75×75\times compared to PVD-DDIM (N=100N=100).

  6. Knowl 6 — Ablation of Distillation Loss Metrics

    data/table

    An ablation study compares the choice of distance metric in flow distillation—comparing standard Euclidean ℓ2\ell_2 loss against Chamfer Distance (CD)—on 1-Nearest Neighbor Accuracy (1-NNA ↓\downarrow) for Airplane, Chair, and Car classes on both PVD-DDIM and PSF.

    Method Dist Airplane Chair Car
    CD EMD CD EMD CD EMD
    PVD-DDIM ℓ2\ell_2 85.5 83.1 82.6 79.3 80.1 77.9
    PVD-DDIM CD 79.9 73.1 72.4 68.2 66.3 65.7
    PSF (ours) ℓ2\ell_2 79.5 73.5 68.4 62.0 67.8 67.0
    PSF (ours) CD 71.1 65.0 59.9 54.3 57.1 56.0

    Using Chamfer Distance during distillation significantly outperforms ℓ2\ell_2 loss across all benchmarks (for instance, improving PSF 1-NNA CD from 79.5 to 71.1 on Airplane and from 68.4 to 59.9 on Chair), demonstrating that a permutation-invariant loss is necessary for distilling unordered 3D point sets.

  7. Knowl 7 — Point Cloud Completion Quality and Latency

    data/table

    Point cloud completion was evaluated using 20-view depth images from ShapeNet Chair, Airplane, and Car categories (following the PVD and GenRe settings). Models take a partial point cloud of 200 points sampled from a single depth image as conditional input to reconstruct complete shapes. Performance is evaluated using Earth Mover's Distance (EMD, multiplied by 10210^2; CD in the benchmark is multiplied by 10310^3).

    Category Model Time (s) ↓\downarrow EMD ↓\downarrow
    Airplane SoftFlow 0.12 1.198
    PointFlow 0.27 1.180
    DPF-Net 0.34 1.105
    PVD (N=1000N=1000) 29.98 1.030
    PSF (ours) 0.04 1.004
    Chair SoftFlow 0.12 3.295
    PointFlow 0.27 3.649
    DPF-Net 0.34 3.320
    PVD (N=1000N=1000) 29.98 2.939
    PSF (ours) 0.04 2.937
    Car SoftFlow 0.12 2.789
    PointFlow 0.27 2.851
    DPF-Net 0.34 2.318
    PVD (N=1000N=1000) 29.98 2.146
    PSF (ours) 0.04 2.194

    PSF achieves completion quality on par with or better than 1000-step PVD while cutting inference latency from ≈30\approx 30 seconds to 0.040.04 seconds.

  8. Knowl 8 — Training-Free Text-Guided Point Cloud Generation

    model/method

    Given a fixed pre-trained Point Straight Flow velocity network vθv_\theta, 3D shapes corresponding to text prompts are generated without additional fine-tuning by directly optimizing the initial Gaussian noise X0X_0.

    Let generator(X0)=X0+vθ(X0,0)\text{generator}(X_0) = X_0 + v_\theta(X_0, 0) be the one-step PSF sampler, and let {Proji}\{\text{Proj}_i\} denote a set of pre-defined camera projections covering uniform viewing angles that render a 3D point cloud into 2D images. The optimization problem is:

    min⁡X0E{Proji}[Sclip(Proji⋅generator(X0), text)]\min_{X_0} \mathbb{E}_{\{\text{Proj}_i\}} \left[ S_{\text{clip}}\left(\text{Proj}_i \cdot \text{generator}(X_0), \, \text{text}\right) \right]

    where SclipS_{\text{clip}} denotes the cosine distance in CLIP embedding space between the rendered 2D view and the input text prompt.

    Optimizing over X0X_0 for 100 iterations requires approximately 12 seconds with the single-step PSF generator, compared to approximately 15 minutes required by a multi-step diffusion baseline such as PVD.

  9. Knowl 9 — Continuous Latent Shape Interpolation via Straight Flows

    empirical result

    Because Point Straight Flow enforces straight deterministic transport paths between Gaussian noise and the data manifold, spherical linear interpolation between two noise samples yields semantically and geometrically continuous shape morphing.

    Let x~0,x~1∼N(0,I)\tilde{x}_0, \tilde{x}_1 \sim \mathcal{N}(0, I) be two random initial noise vectors. An intermediate noise vector x~τ\tilde{x}_\tau is obtained for τ∈[0,1]\tau \in [0, 1] via:

    x~τ=1−τ x~0+τ x~1\tilde{x}_\tau = \sqrt{1 - \tau} \, \tilde{x}_0 + \sqrt{\tau} \, \tilde{x}_1

    Passing x~τ\tilde{x}_\tau into the one-step generator X1′=x~τ+vθ(x~τ,0)X'_1 = \tilde{x}_\tau + v_\theta(\tilde{x}_\tau, 0) yields smooth, continuous geometric transitions between the corresponding 3D shapes. In contrast, multi-step diffusion models (such as PVD) simulated via stochastic differential equations (SDEs) do not maintain this smooth mapping property and generate disjoint intermediate shapes.

Coverage note — None was omitted; all substantive contributions—including the ODE flow formulation, reflow straightness optimization, Chamfer-distance flow distillation, algorithms, empirical evaluations (unconditional generation, completion, text-guided generation, interpolation), and ablations—are covered.

References

  1. 1.Panos Achlioptas, Olga Diamanti, Ioannis Mitliagkas, and Leonidas Guibas. Learning representations and generative models for 3d point clouds. In International conference on machine learning, pages 40–49. PMLR, 2018. 1, 2, 5, 6
  2. 2.Michael S Albergo and Eric Vanden-Eijnden. Building normalizing flows with stochastic interpolants. arXiv preprint arXiv:2209.15571, 2022. 2
  3. 3.Andrew Brock, Theodore Lim, James M Ritchie, and Nick Weston. Generative and discriminative voxel modeling with convolutional neural networks. arXiv preprint arXiv:1608.04236, 2016. 2
  4. 4.Holger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. nuscenes: A multimodal dataset for autonomous driving. In CVPR, 2020. 8, 4
  5. 5.Ruojin Cai, Guandao Yang, Hadar Averbuch-Elor, Zekun Hao, Serge Belongie, Noah Snavely, and Bharath Hariharan. Learning gradient fields for shape generation. In European Conference on Computer Vision, pages 364–381. Springer, 2020. 1, 2, 6
  6. 6.Ricky TQ Chen, Yulia Rubanova, Jesse Bettencourt, and David K Duvenaud. Neural ordinary differential equations. Advances in neural information processing systems, 31, 2018. 2
  7. 7.Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. Advances in Neural Information Processing Systems, 34:8780–8794, 2021. 2
  8. 8.Jun Gao, Tianchang Shen, Zian Wang, Wenzheng Chen, Kangxue Yin, Daiqing Li, Or Litany, Zan Gojcic, and Sanja Fidler. Get3d: A generative model of high quality 3d textured shapes learned from images. In Advances In Neural Information Processing Systems, 2022. 2
  9. 9.Chengyue Gong, Lemeng Wu, and Qiang Liu. How to fill the optimum set? population gradient descent with harmless diversity. arXiv preprint arXiv:2202.08376, 2022. 6
  10. 10.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. Advances in neural information processing systems, 27, 2014. 2
  11. 11.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33:6840–6851, 2020. 1, 2, 4
  12. 12.Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. Video diffusion models. arXiv preprint arXiv:2204.03458, 2022. 2
  13. 13.Hyeongju Kim, Hyeonseung Lee, Woo Hyun Kang, Joun Yeop Lee, and Nam Soo Kim. Softflow: Probabilistic framework for normalizing flow on manifolds. Advances in Neural Information Processing Systems, 33:16388–16397, 2020. 1, 2, 6, 7
  14. 14.Jinwoo Kim, Jaehoon Yoo, Juho Lee, and Seunghoon Hong. Setvae: Learning hierarchical composition for generative modeling of set-structured data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15059–15068, 2021. 1, 2, 5
  15. 15.Diederik P Kingma and Max Welling. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013. 2
  16. 16.Roman Klokov, Edmond Boyer, and Jakob Verbeek. Discrete point flow networks for efficient point cloud generation. In European Conference on Computer Vision, pages 694–710. Springer, 2020. 1, 2, 6, 7
  17. 17.Zhifeng Kong and Wei Ping. On fast sampling of diffusion probabilistic models. arXiv preprint arXiv:2106.00132, 2021. 2
  18. 18.Ruihui Li, Xianzhi Li, Ka-Hei Hui, and Chi-Wing Fu. Sp-gan: Sphere-guided 3d shape generation and manipulation. ACM Transactions on Graphics (TOG), 40(4):1–12, 2021. 2
  19. 19.Angela S. Lin, Lemeng Wu, and Qixing Huang Raymond J. Mooney Rodolfo Corona, Kevin Tai. Generating animated videos of human activities from natural language descriptions. In Proceedings of the Visually Grounded Interaction and Language Workshop at NeurIPS 2018, December 2018. 2
  20. 20.Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. Flow matching for generative modeling. arXiv preprint arXiv:2210.02747, 2022. 1, 2
  21. 21.Qiang Liu. Rectified flow: A marginal preserving approach to optimal transport. arXiv preprint arXiv:2209.14577, 2022. 2
  22. 22.Xingchao Liu, Chengyue Gong, and Qiang Liu. Flow straight and fast: Learning to generate and transfer data with rectified flow. arXiv preprint arXiv:2209.03003, 2022. 1, 2, 3, 4
  23. 23.Xingchao Liu, Chengyue Gong, Lemeng Wu, Shujian Zhang, Hao Su, and Qiang Liu. Fusedream: Training-free text-to-image generation with improved clip+ gan space optimization. arXiv preprint arXiv:2112.01573, 2021. 6
  24. 24.Xingchao Liu, Lemeng Wu, Mao Ye, and Qiang Liu. Let us build bridges: Understanding and extending diffusion generative models. arXiv preprint arXiv:2208.14699, 2022. 2
  25. 25.Zhijian Liu, Haotian Tang, Yujun Lin, and Song Han. Point-voxel cnn for efficient 3d deep learning. Advances in Neural Information Processing Systems, 32, 2019. 4
  26. 26.Eric Luhman and Troy Luhman. Knowledge distillation in iterative generative models for improved sampling speed. arXiv preprint arXiv:2101.02388, 2021. 1, 2, 4
  27. 27.Shitong Luo and Wei Hu. Diffusion probabilistic models for 3d point cloud generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2837–2845, 2021. 1, 2, 6
  28. 28.Zhaoyang Lyu, Zhifeng Kong, Xudong Xu, Liang Pan, and Dahua Lin. A conditional point diffusion-refinement paradigm for 3d point cloud completion. arXiv preprint arXiv:2112.03530, 2021. 2
  29. 29.Oscar Michel, Roi Bar-On, Richard Liu, Sagie Benaim, and Rana Hanocka. Text2mesh: Text-driven neural stylization for meshes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13492–13502, 2022. 6
  30. 30.Alexander Quinn Nichol and Prafulla Dhariwal. Improved denoising diffusion probabilistic models. In International Conference on Machine Learning, pages 8162–8171. PMLR, 2021. 2
  31. 31.George Papamakarios, Eric T Nalisnick, Danilo Jimenez Rezende, Shakir Mohamed, and Balaji Lakshminarayanan. Normalizing flows for probabilistic modeling and inference. J. Mach. Learn. Res., 22(57):1–64, 2021. 2
  32. 32.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022. 2
  33. 33.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10684–10695, 2022. 2
  34. 34.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention, pages 234–241. Springer, 2015. 4
  35. 35.Tim Salimans and Jonathan Ho. Progressive distillation for fast sampling of diffusion models. arXiv preprint arXiv:2202.00512, 2022. 1, 2, 4
  36. 36.Dong Wook Shu, Sung Woo Park, and Junseok Kwon. 3d point cloud generative adversarial network based on tree structured graph convolutions. In Proceedings of the IEEE/CVF international conference on computer vision, pages 3859–3868, 2019. 1
  37. 37.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020. 1, 2, 5, 6
  38. 38.Yang Song and Stefano Ermon. Generative modeling by estimating gradients of the data distribution. Advances in Neural Information Processing Systems, 32, 2019. 2
  39. 39.Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations, 2020. 2
  40. 40.Belinda Tzen and Maxim Raginsky. Theoretical guarantees for sampling and inference in generative models with latent diffusions. In Conference on Learning Theory, pages 3084–3114. PMLR, 2019. 2
  41. 41.Jiajun Wu, Chengkai Zhang, Tianfan Xue, William T Freeman, and Joshua B Tenenbaum. Learning a probabilistic latent space of object shapes via 3d generative-adversarial modeling. In Advances in Neural Information Processing Systems, pages 82–90, 2016. 2
  42. 42.Lemeng Wu, Chengyue Gong, Xingchao Liu, Mao Ye, and Qiang Liu. Diffusion-based molecule generation with informative prior bridges. arXiv preprint arXiv:2209.00865, 2022. 2
  43. 43.Guandao Yang, Xun Huang, Zekun Hao, Ming-Yu Liu, Serge Belongie, and Bharath Hariharan. Pointflow: 3d point cloud generation with continuous normalizing flows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4541–4550, 2019. 1, 2, 5, 6, 7
  44. 44.Ruihan Yang, Prakhar Srivastava, and Stephan Mandt. Diffusion probabilistic modeling for video generation. arXiv preprint arXiv:2203.09481, 2022. 2
  45. 45.Mao Ye, Lemeng Wu, and Qiang Liu. First hitting diffusion models. arXiv preprint arXiv:2209.01170, 2022. 2
  46. 46.Tianwei Yin, Xingyi Zhou, and Philipp Krähenbühl. Multimodal virtual point 3d detection. NeurIPS, 2021. 8
  47. 47.Xiaohui Zeng, Arash Vahdat, Francis Williams, Zan Gojcic, Or Litany, Sanja Fidler, and Karsten Kreis. Lion: Latent point diffusion models for 3d shape generation. In Advances in Neural Information Processing Systems (NeurIPS), 2022. 1, 2
  48. 48.Xiuming Zhang, Zhoutong Zhang, Chengkai Zhang, Josh Tenenbaum, Bill Freeman, and Jiajun Wu. Learning to reconstruct shapes from unseen classes. Advances in neural information processing systems, 31, 2018. 7
  49. 49.Yan Zheng, Lemeng Wu, Xingchao Liu, Zhen Chen, Qiang Liu, and Qixing Huang. Neural volumetric mesh generator. arXiv preprint arXiv:2210.03158, 2022. 2
  50. 50.Linqi Zhou, Yilun Du, and Jiajun Wu. 3d shape generation and completion through point-voxel diffusion. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5826–5835, 2021. 1, 2, 4, 5, 6

Citation

MLA
Wu, L., et al. “Fast Point Cloud Generation with Straight Flows”. arXiv, 2022, http://arxiv.org/abs/2212.01747v1.
APA
Wu, L., Wang, D., Gong, C., Liu, X., Xiong, Y., Ranjan, R., Krishnamoorthi, R., Chandra, V., & Liu, Q. (2022). Fast Point Cloud Generation with Straight Flows. arXiv. http://arxiv.org/abs/2212.01747v1
Chicago
Wu, L., D. Wang, C. Gong, et al. 2022. “Fast Point Cloud Generation with Straight Flows”. arXiv. http://arxiv.org/abs/2212.01747v1.
Harvard
Wu, L. et al. (2022) “Fast Point Cloud Generation with Straight Flows”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2212.01747v1.
Vancouver
1. Wu L, Wang D, Gong C, Liu X, Xiong Y, Ranjan R, Krishnamoorthi R, Chandra V, Liu Q (2022) Fast Point Cloud Generation with Straight Flows. arXiv

BibTeX

@article{wu2022fast,
  title = {Fast Point Cloud Generation with Straight Flows},
  author = {Wu, Lemeng and Wang, Dilin and Gong, Chengyue and Liu, Xingchao and Xiong, Yunyang and Ranjan, Rakesh and Krishnamoorthi, Raghuraman and Chandra, Vikas and Liu, Qiang},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2212.01747v1},
  eprint = {2212.01747}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE