SVGDreamer: Text Guided SVG Generation with Diffusion Model

Ximing XingHaitao ZhouChuang WangJing ZhangDong XuQian Yu

article2024CVPR115 citations

Proposes a text-guided vector graphics generation framework that decomposes visual elements into editable foreground and background layers and uses particle-based score distillation to produce diverse, high-quality SVGs across multiple artistic styles.

Listen

Scalable Vector Graphics (SVGs) are essential for modern graphic design, brand iconography, and digital marketing because they remain sharp at any display resolution and maintain compact file sizes. While recent artificial intelligence systems can generate raster images from text descriptions, converting these text prompts directly into clean, editable vector graphics remains difficult. Existing generative methods typically merge all visual elements into an entangled layer of paths, resulting in outputs that designers cannot easily modify, while frequently suffering from over-smoothed shapes, unnatural color saturation, and limited visual variety.

The article introduces and evaluates SVGDreamer, a novel text-to-SVG framework designed to generate high-quality vector graphics that maintain distinct element editability, visual fidelity, and stylistic diversity. The primary objective is to demonstrate that integrating semantic-driven vectorization with particle-based score distillation solves the structural and visual flaws of existing text-guided vector generation tools.

To achieve this, the authors designed a two-part computational approach. First, a semantic-driven image vectorization (SIVE) module uses cross-attention maps from a pretrained text-to-image diffusion model to separate visual concepts into discrete foreground objects and backgrounds, initializing control points and optimizing them hierarchically with an attention-mask loss. Second, a Vectorized Particle-based Score Distillation (VPSD) method models vector parameters as continuous distributions of control points and colors rather than fixed single values. This process is optimized using low-rank adaptation alongside an aesthetic reward model to accelerate refinement and boost visual appeal. The researchers evaluated the system across six vector art styles—such as iconography, pixel art, and sketches—using 10 distinct prompts per style with 50 outputs generated per prompt, benchmarking against leading baseline methods across standard image quality, diversity, text-alignment, and human preference metrics.

The experimental findings show that SVGDreamer outperforms existing text-to-SVG approaches across all key dimensions. The framework achieved an image diversity score (FID) of 59.13 compared to 100.68 for the previous state of the art, representing an approximate 41% reduction in distribution distance and confirming significantly greater variety. In image fidelity and color balance (PSNR), SVGDreamer reached 14.54 dB, outperforming the leading baseline score of 8.01 dB by over 80% and successfully eliminating oversaturation artifacts. It also recorded top scores in text-image alignment (0.3001 CLIPScore and 0.4623 BLIPScore) as well as overall aesthetic appeal (5.54) and human preference (0.2685). Additionally, the integrated reward feedback mechanism doubled the convergence speed of the optimization process.

These results demonstrate that generative AI can produce practical, modular vector assets that fit directly into professional design workflows. By decomposing scenes into separately editable layers, SVGDreamer enables designers to isolate, adjust, and recombine visual elements rapidly without manual cleanup. This capability significantly lowers asset production timelines and design costs while maintaining high visual quality across diverse branding and illustration tasks.

Organizations seeking to automate asset creation should consider piloting diffusion-guided vector workflows that leverage semantic decomposition. While adoption is recommended, practitioners should note that the system's object separation capabilities remain dependent on the underlying text-to-image diffusion model's ability to isolate complex prompts. Future research should focus on integrating newer base diffusion models and automating the allocation of control points per object layer to further streamline automated graphic production.

Cover for SVGDreamer: Text Guided SVG Generation with Diffusion Model

Abstract

Recently, text-guided scalable vector graphics (SVGs) synthesis has shown promise in domains such as iconography and sketch. However, existing text-to-SVG generation methods lack editability and struggle with visual quality and result diversity. To address these limitations, we propose a novel text-guided vector graphics synthesis method called SVGDreamer. SVGDreamer incorporates a semantic-driven image vectorization (SIVE) process that enables the decomposition of synthesis into foreground objects and background, thereby enhancing editability. Specifically, the SIVE process introduces attention-based primitive control and an attention-mask loss function for effective control and manipulation of individual elements. Additionally, we propose a Vectorized Particle-based Score Distillation (VPSD) approach to address issues of shape over-smoothing, color over-saturation, limited diversity, and slow convergence of the existing text-to-SVG generation methods by modeling SVGs as distributions of control points and colors. Furthermore, VPSD leverages a reward model to re-weight vector particles, which improves aesthetic appeal and accelerates convergence. Extensive experiments are conducted to validate the effectiveness of SVGDreamer, demonstrating its superiority over baseline methods in terms of editability, visual quality, and diversity. Project page: https://ximing.github.io/SVGDreamer-project/

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Vector Graphics Generation
  • 2.2. Text-to-Image Diffusion Model
  • 2.3. Score Distillation Sampling
  • 3. Methodology
  • 3.1. SIVE: Semantic-driven Image Vectorization
  • 3.1.1 Primitive Initialization
  • 3.1.2 Semantic-aware Optimization
  • 3.2. Vectorized Particle-based Score Distillation
  • 3.3. Vector Representation Primitives
  • 4. Experiments
  • 4.1. Qualitative Evaluation
  • 4.2. Quantitative Evaluation
  • 4.3. Ablation Study
  • 4.3.1 SIVE v.s. LIVE [17]
  • 4.3.2 VPSD v.s. LSDS [11, 12] v.s. ASDS [48]
  • 4.4. Applications of SVGDreamer
  • 5. Conclusion
  • References

Citation

MLA
Xing, X., et al. “SVGDreamer: Text Guided SVG Generation with Diffusion Model”. arXiv, 2023, http://arxiv.org/abs/2312.16476v7.
APA
Xing, X., Zhou, H., Wang, C., Zhang, J., Xu, D., & Yu, Q. (2023). SVGDreamer: Text Guided SVG Generation with Diffusion Model. arXiv. http://arxiv.org/abs/2312.16476v7
Chicago
Xing, X., H. Zhou, C. Wang, J. Zhang, D. Xu, and Q. Yu. 2023. “SVGDreamer: Text Guided SVG Generation with Diffusion Model”. arXiv. http://arxiv.org/abs/2312.16476v7.
Harvard
Xing, X. et al. (2023) “SVGDreamer: Text Guided SVG Generation with Diffusion Model”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.16476v7.
Vancouver
1. Xing X, Zhou H, Wang C, Zhang J, Xu D, Yu Q (2023) SVGDreamer: Text Guided SVG Generation with Diffusion Model. arXiv

BibTeX

@article{xing2023svgdreamer,
  title = {SVGDreamer: Text Guided SVG Generation with Diffusion Model},
  author = {Xing, Ximing and Zhou, Haitao and Wang, Chuang and Zhang, Jing and Xu, Dong and Yu, Qian},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.16476v7},
  eprint = {2312.16476}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE