CLIPDraw: Exploring Text-to-Drawing Synthesis through Language-Image Encoders

Kevin FransLisa B. SorosOlaf Witkowski

article2022NeurIPS306 citations

Introduces a training-free text-to-drawing method that uses a pre-trained CLIP model and differentiable rendering to synthesize recognizable, stylistically diverse vector art directly from natural language prompts.

Listen

Generating images directly from natural language prompts traditionally requires training massive neural networks or optimizing pixel grids. These conventional methods often require extensive computing resources, restrict outputs to narrow image distributions, or generate uninterpretable visual textures that fool machine classifiers rather than conveying clear concepts to humans.

The article demonstrates CLIPDraw, an optimization-based algorithm that synthesizes novel drawings from text descriptions without requiring any dedicated model training. By optimizing vector curves rather than raw pixel matrices, the system generates human-recognizable sketches directly aligned with text prompts.

The researchers designed CLIPDraw using a frozen, pre-trained dual language-image encoder (CLIP) to evaluate how closely a generated visual matches a text prompt. Instead of modifying pixels, the method initializes a fixed number of transparent, colored Bézier vector strokes on a white canvas and refines their positions, thicknesses, and colors using gradient descent over 250 iterations. Crucially, the process applies geometric distortions and crops to the rendered strokes during each step to ensure the final drawing remains robust and recognizable. The article evaluated this approach across varied prompts, comparing it against pixel optimization, generative models such as BigGAN and VQGAN, and unaugmented baselines across multiple random trials.

The evaluation revealed several key findings regarding performance, speed, and visual behavior. First, CLIPDraw reliably synthesizes recognizable sketches within one to two minutes on standard hardware, outperforming direct pixel optimization in visual structure and matching the prompt-alignment scores of more complex generative models. Second, enforcing a vector-stroke constraint inherently biases outputs toward simple, salient human shapes, scaling smoothly from abstract line art at low stroke counts (such as 16 strokes) to detailed, shaded compositions at higher counts (such as 256 strokes). Third, the method demonstrates high stylistic versatility, adapting scene geometry and visual styles—from watercolors to wireframes—purely through text adjectives without requiring separate style-transfer networks. Fourth, the system exhibits creative problem-solving behaviors, such as incorporating literal words or depicting multiple interpretations of ambiguous concepts (for instance, showing both runners and burgers for "Fast Food"). Lastly, image augmentation is essential; running optimization without it produced high numerical match scores but resulted in visual nonsense to human viewers.

These findings indicate that lightweight optimization over constrained geometric primitives can bypass the high computational costs and data overhead of training large generative image models. For product teams and creative industries, this approach enables rapid, low-cost conceptual prototyping and AI-assisted art creation. Furthermore, because it does not depend on a rigid image generator, the system provides an agile, transparent testbed to inspect how large language-vision models associate abstract words with visual concepts.

Decision-makers and practitioners should consider adopting vector-based optimization when building low-overhead creative tools or exploratory interfaces. To further enhance control, teams should investigate refined steering mechanisms, as initial attempts to use negative text prompts to suppress unwanted elements showed inconsistent quality improvements. Organizations should also establish human oversight when applying these tools commercially, as the synthesized imagery inherently inherits the cultural and social biases embedded in the underlying pre-trained foundation model.

The primary limitations of CLIPDraw involve its difficulty with high-resolution photorealism and its lack of fine-grained spatial control, such as precisely placing an object in a specific region of the canvas. The findings are highly reliable within the bounded scope of vector sketches and qualitative conceptual synthesis, though readers should note that the system is optimized for stylized visual ideation rather than exact graphical layouts.

arXiv: 2106.14843
Cover for CLIPDraw: Exploring Text-to-Drawing Synthesis through Language-Image Encoders

Abstract

CLIPDraw is an algorithm that synthesizes novel drawings from natural language input. It does not require any additional training; rather, a pre-trained CLIP language-image encoder is used as a metric for maximizing similarity between the given description and a generated drawing. Crucially, CLIPDraw operates over vector strokes rather than pixel images, which biases drawings towards simpler human-recognizable shapes. Results compare CLIPDraw with other synthesis-through-optimization methods, as well as highlight various interesting behaviors of CLIPDraw, such as satisfying ambiguous text in multiple ways, reliably producing drawings in diverse styles, and scaling from simple to complex visual representations as stroke count increases.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 4 Results
  • 4.1 How does CLIPDraw compare to other synthesis-through-optimization methods?
  • 4.2 What kinds of visual techniques does CLIPDraw use to satisfy the textual description?
  • 4.3 Can CLIPDraw reliably produce drawings in different styles?
  • 4.4 How does the stroke count affect what drawings CLIPDraw produces?
  • 4.5 What happens if abstract words are given as a description prompt?
  • 4.6 Can synthesized drawings be fine-tuned via additional negative prompts?
  • 5 Discussion
  • 5.1 Limitations
  • 5.2 Ethics and Social Biases
  • 5.3 Future Work
  • References

Knowls

  1. Knowl 1 — CLIPDraw Framework for Vector Drawing Synthesis

    model/method

    CLIPDraw is a synthesis-through-optimization method that generates vector drawings from natural language descriptions without requiring generative model training or vector datasets. An image is represented as a fixed number NN of differentiable Bézier curves over a white background canvas. Each curve is parameterized by 33 to 55 2D control points, a stroke thickness parameter, and an RGBA color vector.

    To synthesize a drawing matching a text prompt, the curve parameters are optimized directly at test time via gradient descent. In each iteration, a differentiable vector graphics renderer (DiffVG) rasterizes the curves into a pixel image. To prevent the optimization from converging to human-unrecognizable adversarial artifacts, the rendered image is replicated DD times and passed through stochastic differentiable augmentations (a sequence of random perspective transformation and random resized crop). The augmented views and the target text prompt are mapped into a shared 512-dimensional embedding space using a frozen pre-trained CLIP (Contrastive Language-Image Pre-training) model, and gradient updates backpropagate through the CLIP image encoder, augmentations, and differentiable rasterizer to update the curve parameters.

  2. Knowl 2 — CLIPDraw Optimization Loop

    algorithm

    The CLIPDraw algorithm generates a vector drawing matching a textual prompt descdesc over II gradient descent iterations.

    Input: Description phrase descdesc, iteration count II, curve count NN, augmentation batch size DD, pre-trained CLIP model with text encoder CLIPtext\text{CLIP}_{\text{text}} and image encoder CLIPimage\text{CLIP}_{\text{image}}, differentiable vector renderer DiffRender\text{DiffRender}, stochastic image augmentation module Augment\text{Augment}.
    Output: Optimized set of Bézier curves CurvesCurves.
    EncPhr←CLIPtext(desc)EncPhr \leftarrow \text{CLIP}_{\text{text}}(desc)
    Curves←InitializeRandomCurves(N)Curves \leftarrow \text{InitializeRandomCurves}(N)
    for i←0i \leftarrow 0 to I−1I - 1 do
        Pixels←DiffRender(Curves)Pixels \leftarrow \text{DiffRender}(Curves)
        AugBatch←Augment(Pixels,D)AugBatch \leftarrow \text{Augment}(Pixels, D)
        EncImg←CLIPimage(AugBatch)EncImg \leftarrow \text{CLIP}_{\text{image}}(AugBatch)
        Loss←−∑d=1DCosineSim(EncPhr,EncImgd)Loss \leftarrow -\sum_{d=1}^{D} \text{CosineSim}(EncPhr, EncImg_d)
        Curves←GradientUpdate(Curves,∇CurvesLoss)Curves \leftarrow \text{GradientUpdate}(Curves, \nabla_{Curves} Loss)
    end for
    return CurvesCurves

    In standard execution, N=256N = 256 strokes, D=8D = 8 augmented copies per step, and optimization runs for I=250I = 250 gradient descent iterations, taking approximately 1 to 2 minutes on a standard GPU.

  3. Knowl 3 — CLIPDraw Text-Matching and Negative-Prompt Optimization Objectives

    equation

    Let eprompt=CLIPtext(desc)∈R512e_{\text{prompt}} = \text{CLIP}_{\text{text}}(desc) \in \mathbb{R}^{512} be the normalized CLIP text embedding of the input description, and let eimg(d)=CLIPimage(Ad(DiffRender(C)))∈R512e_{\text{img}}^{(d)} = \text{CLIP}_{\text{image}}(A_d(\text{DiffRender}(C))) \in \mathbb{R}^{512} be the normalized CLIP image embedding of the dd-th augmented rasterized view of curve parameters CC, where AdA_d represents a composition of random perspective transformations and random resized crops for d∈{1,…,D}d \in \{1, \dots, D\}.

    The primary loss function minimized with respect to CC is the negative average cosine similarity: L(C)=−1D∑d=1Deprompt⋅eimg(d)∥eprompt∥2∥eimg(d)∥2\mathcal{L}(C) = -\frac{1}{D} \sum_{d=1}^D \frac{e_{\text{prompt}} \cdot e_{\text{img}}^{(d)}}{\|e_{\text{prompt}}\|_2 \|e_{\text{img}}^{(d)}\|_2}

    When using a set of MM negative prompts with text encodings eneg(m)=CLIPtext(neg_descm)e_{\text{neg}}^{(m)} = \text{CLIP}_{\text{text}}(neg\_desc_m) and a penalty weighting factor λ\lambda (e.g., λ=0.3\lambda = 0.3), the objective penalizes similarity to the negative prompts: Ltotal(C)=−1D∑d=1D(eprompt⋅eimg(d)∥eprompt∥2∥eimg(d)∥2−λ∑m=1Meneg(m)⋅eimg(d)∥eneg(m)∥2∥eimg(d)∥2)\mathcal{L}_{\text{total}}(C) = -\frac{1}{D} \sum_{d=1}^D \left( \frac{e_{\text{prompt}} \cdot e_{\text{img}}^{(d)}}{\|e_{\text{prompt}}\|_2 \|e_{\text{img}}^{(d)}\|_2} - \lambda \sum_{m=1}^M \frac{e_{\text{neg}}^{(m)} \cdot e_{\text{img}}^{(d)}}{\|e_{\text{neg}}^{(m)}\|_2 \|e_{\text{img}}^{(d)}\|_2} \right)

  4. Knowl 4 — Optimization-Based Text-to-Image Baseline Configurations

    experimental setup

    To evaluate synthesis quality, CLIPDraw is compared against alternative optimization-based synthesis methods that use the same pre-trained CLIP objective for 250 gradient descent steps:

    1. CLIPDraw: Optimizes N=256N=256 RGBA Bézier curves with D=8D=8 perspective/crop augmentations per step.
    2. Pixel Optimization: Directly optimizes a 224×224×3224 \times 224 \times 3 RGB pixel matrix under identical image augmentations and CLIP objective.
    3. BigGAN Optimization: Optimizes the latent vector zz of a frozen pre-trained BigGAN generator.
    4. VQGAN Optimization: Optimizes representations through sampling the latent codebook of a pre-trained VQGAN generator.
    5. CLIPDraw (No Augment): Identical to CLIPDraw, but optimizes the curve parameters directly against the unaugmented rendered pixel image (D=1D=1, no affine/crop distortions).

    Evaluation is conducted over 10 random seeds per prompt across diverse textual descriptions.

  5. Knowl 5 — Quantitative Evaluation of CLIP Cosine Similarities Across Synthesis Baselines

    data/table

    The quality of prompt adherence is measured by computing the average CLIP cosine similarity between the textual prompt and 64 random augmentations of the generated output across 10 random initialization seeds after 250 gradient steps (mean ±\pm standard deviation):

    Prompt CLIPDraw Pixel Opt. BigGAN Opt VQGAN Opt CLIPDraw (No Aug)
    A drawing of a cat. .376 ±\pm .005 .385 ±\pm .009 .325 ±\pm .014 .377 ±\pm .004 .240 ±\pm .015
    A paint. of a sunset. .390 ±\pm .004 .215 ±\pm .010 .379 ±\pm .003 .379 ±\pm .003 .325 ±\pm .015
    Underwater. .413 ±\pm .006 .385 ±\pm .009 .327 ±\pm .013 .358 ±\pm .006 .247 ±\pm .004
    Sheep wearing a top hat. .434 ±\pm .010 .434 ±\pm .010 .321 ±\pm .016 .411 ±\pm .009 .215 ±\pm .011
    A 3D rend. of a temple. .467 ±\pm .005 .385 ±\pm .009 .311 ±\pm .028 .430 ±\pm .008 .240 ±\pm .009
    Watercol. of a firetruck. .507 ±\pm .007 .385 ±\pm .009 .318 ±\pm .014 .457 ±\pm .007 .233 ±\pm .007
    Third Eye. .436 ±\pm .010 .385 ±\pm .009 .323 ±\pm .015 .368 ±\pm .005 .251 ±\pm .013

    CLIPDraw matches or exceeds the semantic similarity of GAN-constrained baselines across multi-attribute and structured prompts. While Pixel Optimization achieves comparable raw cosine scores on certain simple prompts, its outputs consist of noisy high-frequency textures lacking coherent shapes. CLIPDraw without augmentation achieves drastically lower test-time similarity (0.215−0.3250.215 - 0.325) because it overfits to unaugmented views with human-unrecognizable adversarial strokes.

  6. Knowl 6 — Stroke Count as a Simplicity-to-Complexity Visual Bias

    empirical result

    Varying the number of vector strokes NN in CLIPDraw acts as a direct control parameter balancing abstract simplicity and visual realism:

    • Low stroke counts (N=16N = 16 to 3232): The representation constraint forces the model to discard fine details and textures, producing minimalistic cartoon sketches and geometric outlines that isolate the essential identifiable structure of the prompt (e.g., drawing the Eiffel Tower using only a few straight line segments).
    • High stroke counts (N=128N = 128 to 512512): The optimization utilizes additional curve capacity to introduce complex lighting, shading, depth cues, and rich background contexts (e.g., skies, surrounding structures, and multi-layered colors).
  7. Knowl 7 — Textual Control of Artistic Style and Spatial Representation

    empirical result

    Because CLIPDraw is not constrained to the natural image manifold of a pre-trained generative network (such as BigGAN or VQGAN), specifying style-descriptive adjectives in the prompt alters both textural surface appearance and underlying structural geometry:

    • Specifying "a drawing of a cat" produces flat, 2D outline sketches.
    • Specifying "a 3D wireframe model" or "a cat as 3D rendered in Unreal Engine" alters the spatial composition to depict volumetric perspective, depth-based blurring, directional lighting, and shadows.
    • Specifying artistic mediums (e.g., "watercolor painting", "Japanese woodblock print") biases stroke colors, gradients, and transparency toward palette distributions characteristic of those media.
  8. Knowl 8 — Emergence of Typographic Inscription and Semantic Disambiguation

    empirical result

    CLIPDraw exhibits distinct emergent behaviors driven by the joint language-vision representations of CLIP:

    1. Text and Letter Inscription: The optimizer frequently synthesizes literal text, words, or character-like glyphs directly into the drawing canvas corresponding to keywords in the prompt (e.g., spelling "Yeti" in "Yeti taking a selfie" or writing "third" and "eye" in "Third Eye").
    2. Multi-Interpretation of Polysemous Words: When presented with semantically ambiguous prompts, CLIPDraw often illustrates multiple meanings simultaneously within a single composition. For example, for the prompt "Fast Food", CLIPDraw generates both hamburgers / fast-food logos and human runners in a footrace, with CLIP zero-shot classification confirming high activation for both "hamburger" and "jogging".
  9. Knowl 9 — Symbolic and Metaphorical Grounding of Abstract Concepts

    empirical result

    When prompted with abstract words that lack concrete physical referents, CLIPDraw synthesizes drawings composed of culturally and symbolically associated imagery captured in CLIP's training corpus:

    • "Happiness" produces drawings with smiling faces, bright colors, and fireworks.
    • "Translation" depicts bilingual mixtures of English and Japanese-like typographic glyphs.
    • "Enlightenment" synthesizes a seated monk-like silhouette.
    • "Self" generates a human body containing multiple distinct heads, capturing the psychological concept of multi-faceted personal identity.
    • "What do you look like, CLIPDraw?" generates a smiling face alongside strokes spelling the text "CLIPDRAW".
  10. Knowl 10 — Limitations in Spatial Layout Steering, Resolution, and Negative Prompt Stability

    limitation

    CLIPDraw has three primary operational limitations:

    1. Lack of Fine-Grained Spatial Steering: Because the loss operates globally on holistic CLIP embeddings, fine spatial adjustments (such as relocating a specific object to another side of the canvas) cannot be reliably commanded via textual descriptions alone.
    2. Resolution and Realism Ceiling: Constraining synthesis purely to vector Bézier curves prevents the generation of photorealistic fine textures and high-resolution details compared to high-capacity generative models.
    3. Inconsistent Negative Prompting: While subtracting negative prompt similarities can occasionally reduce unwanted features (e.g., suppressing text by penalizing "Words and text"), negative prompts frequently exhibit negligible effects, and universal quality-improving negative prompts (e.g., "a low-quality drawing") fail to reliably improve drawing aesthetics.

Coverage note — None was omitted; all primary methodology, algorithmic formulations, experimental comparisons, qualitative behaviors, and stated limitations are included.

References

  1. 1.Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al. (2021). On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258.
  2. 2.Chen, M., Radford, A., Child, R., Wu, J., Jun, H., Luan, D., and Sutskever, I. (2020). Generative pretraining from pixels. In International Conference on Machine Learning, pages 1691–1703. PMLR.
  3. 3.Crowson, K., Biderman, S., Kornis, D., Stander, D., Hallahan, E., Castricato, L., and Raff, E. (2022). Vqgan-clip: Open domain image generation and editing with natural language guidance. arXiv preprint arXiv:2204.08583.
  4. 4.Erhan, D., Bengio, Y., Courville, A., and Vincent, P. (2009). Visualizing higher-layer features of a deep network. University of Montreal, 1341(3):1.
  5. 5.Esser, P., Rombach, R., and Ommer, B. (2021). Taming transformers for high-resolution image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12873–12883.
  6. 6.Fernando, C., Eslami, S., Alayrac, J.-B., Mirowski, P., Banarse, D., and Osindero, S. (2021). Generative art using neural visual grammars and dual encoders. arXiv preprint arXiv:2105.00162.
  7. 7.Frolov, S., Hinz, T., Raue, F., Hees, J., and Dengel, A. (2021). Adversarial text-to-image synthesis: A review. arXiv preprint arXiv:2101.09983.
  8. 8.Galatolo, F. A., Cimino, M. G., and Vaglini, G. (2021). Generating images from caption and vice versa via clip-guided generative latent space search. arXiv preprint arXiv:2102.01645.
  9. 9.Gatys, L. A., Ecker, A. S., and Bethge, M. (2015). A neural algorithm of artistic style. arXiv preprint arXiv:1508.06576.
  10. 10.Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014a). Generative adversarial networks. arXiv preprint arXiv:1406.2661.
  11. 11.Goodfellow, I. J., Shlens, J., and Szegedy, C. (2014b). Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572.
  12. 12.Kotovenko, D., Wright, M., Heimbrecht, A., and Ommer, B. (2021). Rethinking style transfer: From pixels to parameterized brushstrokes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12196–12205.
  13. 13.Li, T.-M., Lukáč, M., Gharbi, M., and Ragan-Kelley, J. (2020). Differentiable vector graphics rasterization for editing and learning. ACM Transactions on Graphics (TOG), 39(6):1–15.
  14. 14.Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L. (2014). Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer.
  15. 15.Mirowski, P., Banarse, D., Malinowski, M., Osindero, S., and Fernando, C. (2022). Clip-clop: Clip-guided collage and photomontage. arXiv preprint arXiv:2205.03146.
  16. 16.Mirza, M. and Osindero, S. (2014). Conditional generative adversarial nets. arXiv preprint arXiv:1411.1784.
  17. 17.Mordvintsev, A., Olah, C., and Tyka, M. (2015). Inceptionism: Going deeper into neural networks.
  18. 18.Murdock, R. (2021). The big sleep: Bigganxclip.
  19. 19.Nguyen, A., Clune, J., Bengio, Y., Dosovitskiy, A., and Yosinski, J. (2017). Plug & play generative networks: Conditional iterative generation of images in latent space. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4467–4477.
  20. 20.Nguyen, A., Dosovitskiy, A., Yosinski, J., Brox, T., and Clune, J. (2016). Synthesizing the preferred inputs for neurons in neural networks via deep generator networks. arXiv preprint arXiv:1605.09304.
  21. 21.Nguyen, A., Yosinski, J., and Clune, J. (2015). Deep neural networks are easily fooled: High confidence predictions for unrecognizable images. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 427–436.
  22. 22.Nilsback, M.-E. and Zisserman, A. (2008). Automated flower classification over a large number of classes. In 2008 Sixth Indian Conference on Computer Vision, Graphics & Image Processing, pages 722–729. IEEE.
  23. 23.Oord, A. v. d., Vinyals, O., and Kavukcuoglu, K. (2017). Neural discrete representation learning. arXiv preprint arXiv:1711.00937.
  24. 24.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. (2021). Learning transferable visual models from natural language supervision. arXiv preprint arXiv:2103.00020.
  25. 25.Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I. (2021). Zero-shot text-to-image generation. arXiv preprint arXiv:2102.12092.
  26. 26.Reddy, P., Gharbi, M., Lukac, M., and Mitra, N. J. (2021). Im2vec: Synthesizing vector graphics without vector supervision. arXiv preprint arXiv:2102.02798.
  27. 27.Reed, S., Akata, Z., Yan, X., Logeswaran, L., Schiele, B., and Lee, H. (2016). Generative adversarial text to image synthesis. In International Conference on Machine Learning, pages 1060–1069. PMLR.
  28. 28.Schaldenbrand, P., Liu, Z., and Oh, J. (2021). Styleclipdraw: Coupling content and style in text-to-drawing synthesis. arXiv preprint arXiv:2111.03133.
  29. 29.Shen, I.-C. and Chen, B.-Y. (2021). Clipgen: A deep generative model for clipart vectorization and synthesis. IEEE Transactions on Visualization and Computer Graphics, pages 1–1.
  30. 30.Tian, Y. and Ha, D. (2022). Modern evolution strategies for creativity: Fitting concrete images and abstract concepts. In International Conference on Computational Intelligence in Music, Sound, Art and Design (Part of EvoStar), pages 275–291. Springer.
  31. 31.Vinker, Y., Pajouheshgar, E., Bo, J. Y., Bachmann, R. C., Bermano, A. H., Cohen-Or, D., Zamir, A., and Shamir, A. (2022). Clipasso: Semantically-aware object sketching. arXiv preprint arXiv:2202.05822.
  32. 32.Wah, C., Branson, S., Welinder, P., Perona, P., and Belongie, S. (2011). The caltech-ucsd birds-200-2011 dataset.

Citation

MLA
Frans, K., et al. “CLIPDraw: Exploring Text-to-Drawing Synthesis Through Language-Image Encoders”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 5207–18, https://proceedings.neurips.cc/paper_files/paper/2022/file/21f76686538a5f06dc431efea5f475f5-Paper-Conference.pdf.
APA
Frans, K., Soros, L., & Witkowski, O. (2022). CLIPDraw: Exploring Text-to-Drawing Synthesis through Language-Image Encoders. Advances in Neural Information Processing Systems, 35, 5207–5218. https://proceedings.neurips.cc/paper_files/paper/2022/file/21f76686538a5f06dc431efea5f475f5-Paper-Conference.pdf
Chicago
Frans, K., L. Soros, and O. Witkowski. 2022. “CLIPDraw: Exploring Text-to-Drawing Synthesis Through Language-Image Encoders”. Advances in Neural Information Processing Systems 35: 5207–18. https://proceedings.neurips.cc/paper_files/paper/2022/file/21f76686538a5f06dc431efea5f475f5-Paper-Conference.pdf.
Harvard
Frans, K., Soros, L. and Witkowski, O. (2022) “CLIPDraw: Exploring Text-to-Drawing Synthesis through Language-Image Encoders”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 5207–5218. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/21f76686538a5f06dc431efea5f475f5-Paper-Conference.pdf.
Vancouver
1. Frans K, Soros L, Witkowski O (2022) CLIPDraw: Exploring Text-to-Drawing Synthesis through Language-Image Encoders. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 5207–5218

BibTeX

@inproceedings{frans2022clipdraw,
  title = {CLIPDraw: Exploring Text-to-Drawing Synthesis through Language-Image Encoders},
  author = {Frans, Kevin and Soros, Lisa and Witkowski, Olaf},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {5207-5218},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/21f76686538a5f06dc431efea5f475f5-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors