Text-to-CAD Generation Through Infusing Visual Feedback in Large Language Models

Ruiyu WangYu YuanShizhao SunJiang Bian

article2025ICML53 citations

Introduces CADFusion, an alternating training framework that enables large language models to generate precise 3D CAD models from text by combining parametric sequence supervision with rendered visual feedback.

Listen

Computer-Aided Design (CAD) is essential for modern engineering and manufacturing, yet creating 3D CAD models remains a manual, time-intensive process requiring specialized technical expertise. Automated "Text-to-CAD" generation—converting natural language descriptions directly into executable parametric design sequences—offers a way to accelerate prototyping and democratize 3D modeling. However, CAD models are inherently multimodal: they exist both as sequential geometric commands and as rendered visual objects. Furthermore, because multiple distinct command sequences can produce the exact same 3D shape, training systems solely on ground-truth command sequences leads models to memorize specific sequence patterns while failing to capture global visual geometry.

The article introduces CADFusion, an artificial intelligence framework designed to evaluate and demonstrate how combining sequential code learning with automated visual feedback enables large language models to generate accurate, high-quality CAD parametric models from plain text descriptions.

The researchers utilized a fine-tuned, 8-billion-parameter open-source language model as the core engine, pairing it with a Sketch-and-Extrude representation where CAD operations are expressed as text tokens. Training alternated between two recurring stages: a sequential learning phase fine-tuning the model on 20,000 paired text-CAD examples refined by human annotators, and a visual feedback phase. To bypass the non-differentiable barrier of CAD rendering engines, the authors framed visual learning as a preference optimization task. Rendered 3D outputs were automatically scored across shape quality, component quantity, and spatial distribution using vision-language models, generating preference rankings without relying on expensive human evaluations. The framework was quantitatively evaluated against general-purpose models (GPT-4o) and specialized baseline architectures across geometric fidelity, rendering validity, visual alignment, and human expert rankings.

The experimental findings show that CADFusion substantially outperforms existing approaches across key operational metrics. First, CADFusion achieved a visual evaluation score of 8.96 out of 10 and an average human preference rank of 1.86, outperforming both general-purpose GPT-4o (score 5.13; rank 3.22) and specialized prior systems (rank 2.97). Second, the model maintained high sequence validity, rendering successfully with an invalidity ratio of only 6.20%, compared to a 74.26% failure rate for GPT-4o. Third, in geometric accuracy, CADFusion demonstrated superior point-cloud alignment with a Chamfer Distance of 19.89 and a 90.40% coverage rate, whereas prior methods struggled with complex geometries. Finally, ablation experiments confirmed that alternating between sequential training and visual feedback was essential; removing visual feedback reduced visual quality scores from 8.96 to 7.69, while visual training without alternating sequential learning caused invalidity rates to surge to 88.87% due to syntax degradation.

These results demonstrate that infusing visual preference signals into language models overcomes the core limitations of command-only text-to-CAD systems. In practice, this enables engineering organizations to rapidly generate valid, modifiable CAD assets from concise, non-expert text prompts rather than requiring tedious step-by-step drafting instructions. Because the model supports non-deterministic sampling, designers can quickly produce diverse design variations with adjustable parameters, lowering development cycle times and design iteration costs.

Organizations evaluating automated CAD workflows should consider pilot implementations using alternating sequence-and-preference optimization frameworks rather than relying on out-of-the-box language models or pure sequence generators. Development teams should prioritize curated, human-refined text prompts during initial training, as scaling unrefined synthetic datasets yielded negligible visual improvements. Future engineering initiatives must focus on incorporating multi-view visual feedback pipelines and expanding training datasets to include more intricate geometric assemblies.

These findings carry moderate-to-high confidence for standardized geometric structures and parametric primitives. However, decision-makers should exercise caution regarding current technical boundaries: the visual feedback pipeline currently relies on single-view image evaluations due to vision model constraints, and the system still struggles with highly complex spatial reasoning tasks, such as generating text-shaped geometry or coordinating more than eight distinct sub-components.

No sufficiently relevant recommendations were found.

No sufficiently relevant recommendations were found.

Cover for Text-to-CAD Generation Through Infusing Visual Feedback in Large Language Models

Abstract

Creating Computer-Aided Design (CAD) models requires significant expertise and effort. Text-to-CAD, which converts textual descriptions into CAD parametric sequences, is crucial in streamlining this process. Recent studies have utilized ground-truth parametric sequences, known as sequential signals, as supervision to achieve this goal. However, CAD models are inherently multimodal, comprising parametric sequences and corresponding rendered visual objects. Besides,the rendering process from parametric sequences to visual objects is many-to-one. Therefore, both sequential and visual signals are critical for effective training. In this work, we introduce CADFusion, a framework that uses Large Language Models (LLMs) as the backbone and alternates between two training stages: the sequential learning (SL) stage and the visual feedback (VF) stage. In the SL stage, we train LLMs using ground-truth parametric sequences, enabling the generation of logically coherent parametric sequences. In the VF stage, we reward parametric sequences that render into visually preferred objects and penalize those that do not, allowing LLMs to learn how rendered visual objects are perceived and evaluated. These two stages alternate throughout the training, ensuring balanced learning and preserving benefits of both signals. Experiments demonstrate that CADFusion significantly improves performance, both qualitatively and quantitatively.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 Method
  • 3.1 Approach Overview
  • 3.2 Sequential Learning Stage
  • 3.3 Visual Feedback Stage
  • 3.4 Alternate Training
  • 4 Experiments
  • 4.1 Setups
  • 4.2 Main Results
  • 4.3 Ablation Studies
  • 5 Limitation
  • 6 Conclusion
  • References
  • A User Guidelines for Prompting
  • B Additional Dataset Construction Detail
  • B.1 Converting Raw Data into Strings
  • B.2 Generating Textual Instructions
  • B.3 Dataset Construction
  • C Additional Training Detail
  • C.1 Sequential Learning
  • C.2 Visual Feedback
  • D Additional Experimental Results
  • D.1 LVM Evaluation Setups
  • D.2 Human Evaluation Setups
  • D.3 GPT Baselines
  • D.4 Additional Statements on Text2CAD Results
  • D.5 Additional Quantative Results
  • D.6 Additional Qualitative Results
  • D.7 Text to Multiple CAD Figures
  • D.8 Failure Cases
  • E Change Log

Knowls

  1. Knowl 1 — CADFusion combines sequential supervision with rendered-visual preferences

    model/method

    CADFusion generates a CAD parametric sequence from a text description using a pretrained large language model (LLM), then renders the sequence into a visual object. Training uses two complementary signals: ground-truth sequences teach the model valid construction structure and operations, while preferences between rendered objects teach it which outputs better match the description. The framework alternates sequential learning and visual-feedback training so that visual optimization does not replace sequence-format learning, and sequence-only optimization does not neglect rendered appearance. This design addresses the fact that different parametric sequences can render to the same visual object.

  2. Knowl 2 — Visual feedback is learned with direct preference optimization

    equation

    Let xx be a text description, ywy_w a CAD sequence whose rendered object is preferred, and yly_l a sequence whose rendered object is less preferred. Let pθ(y∣x)p_\theta(y\mid x) be the probability of sequence yy under the current LLM, and let pref(y∣x)p_{\mathrm{ref}}(y\mid x) be its probability under the fixed reference LLM from the preceding sequential-learning stage. For preference triples (x,yw,yl)(x,y_w,y_l) drawn from the visual-feedback dataset DVF\mathcal{D}_{\mathrm{VF}}, CADFusion minimizes the DPO objective

    LVF=−E(x,yw,yl)∼DVF[log⁡σ(β(log⁡pθ(yw∣x)pref(yw∣x)−log⁡pθ(yl∣x)pref(yl∣x)))],\mathcal{L}_{\mathrm{VF}}=-\mathbb{E}_{(x,y_w,y_l)\sim\mathcal{D}_{\mathrm{VF}}}\left[\log\sigma\left(\beta\left(\log\frac{p_\theta(y_w\mid x)}{p_{\mathrm{ref}}(y_w\mid x)}-\log\frac{p_\theta(y_l\mid x)}{p_{\mathrm{ref}}(y_l\mid x)}\right)\right)\right],

    where σ\sigma is the logistic sigmoid and β\beta is a scaling factor. The preferred and less-preferred labels come from comparing rendered objects, so optimization does not require differentiating through the non-differentiable CAD-rendering process.

  3. Knowl 3 — Alternating training preserves both sequence validity and visual quality

    model/method

    CADFusion first trains on ground-truth CAD sequences, then repeats blocks consisting of visual-feedback training followed by sequential learning. The initial sequence-learning phase gives the LLM the ability to produce coherent, renderable command sequences. In the reported configuration, the initial phase lasts 40 epochs; five alternating rounds follow, each with five visual-feedback epochs and one sequential-learning epoch. The backbone is LLaMA-3-8B-Instruct with a 1024-token maximum length, fine-tuned using LoRA with rank 3232 and α=32\alpha=32; the initial sequential-learning rate is 1×10−41\times10^{-4} with AdamW. Training uses four NVIDIA A6000-48GB GPUs with distributed data parallelism. The authors report that visual-feedback-only training can damage sequence formatting, while sequence-only training can weaken rendered-object quality; alternation is intended to limit both forms of degradation.

  4. Knowl 4 — Sequential learning fine-tunes an LLM to emit text-serialized Sketch-and-Extrude CAD

    model/method

    The sequential-learning stage fine-tunes a pretrained LLM on paired descriptions and CAD parametric sequences. CAD is represented in a text-based Sketch-and-Extrude format: sketches contain lines, arcs, and circles organized into loops and faces, followed by extrusion commands such as add, cut, or intersect and their parameters. Structural markers delimit curves, loops, faces, sketches, and extrusions, and the sequence is processed with the LLM tokenizer rather than a separate learned codebook. For a training pair (x,y)(x,y), where xx is the description and y=(y1,…,yT)y=(y_1,\ldots,y_T) is a sequence of TT CAD tokens, the model minimizes average autoregressive token cross-entropy:

    LSL=−E(x,y)∼DSL[1T∑t=1Tlog⁡pθ(yt∣x,y<t)].\mathcal{L}_{\mathrm{SL}}=-\mathbb{E}_{(x,y)\sim\mathcal{D}_{\mathrm{SL}}}\left[\frac{1}{T}\sum_{t=1}^{T}\log p_\theta(y_t\mid x,y_{<t})\right].

    Here pθp_\theta is the current model’s token probability and y<ty_{<t} denotes the preceding ground-truth tokens. This stage supplies direct supervision for valid sequence structure and parametric operations.

  5. Knowl 5 — A vision-language model constructs visual-preference pairs automatically

    model/method

    To create preference data, CADFusion samples multiple CAD sequences from the model after sequential learning for each text prompt, renders the sequences, and asks a large vision-language model (LVM) to score the resulting CAD images. The scoring instruction evaluates three aspects: whether the overall shape is regular and natural, whether the number of components matches the description, and whether components are arranged without collisions or excessive spacing. The higher-scoring rendered object is labeled preferred and the lower-scoring one less preferred; their corresponding CAD sequences form a DPO training pair. The reported scoring model is LLaVA-OneVision-Qwen2-7B, with scores on a 0–10 scale. Each visual-feedback round uses 1,000 prompts and five samples per prompt, filtering invalid or low-quality outputs to obtain approximately 1,500 preference pairs.

  6. Knowl 6 — The training data pairs human-refined captions with serialized CAD sequences

    data/table

    The sequential-learning dataset contains 20,000 pairs of text instructions and CAD parametric sequences, sourced from the DeepCAD data processed by Xu and colleagues. For each CAD model, an LVM first drafts a caption from a rendered image; human annotators then correct inaccuracies and make the description concise. Captions that are too difficult to describe may be excluded. The paired CAD construction history is converted into the text-based Sketch-and-Extrude sequence format. The dataset is split into training, validation, and test sets in a 90:5:5 ratio. This dataset supplies the ground-truth sequential signal, while the separately generated rendered-object preference pairs supply the visual signal.

  7. Knowl 7 — CADFusion leads the reported evaluation, especially on visual correspondence

    empirical result

    On the test set, CADFusion reports the following metric vectors, in this order: F1-Sketch, F1-Extrusion, Chamfer Distance (CD), Coverage (COV), Minimum Matching Distance (MMD), Jensen–Shannon Divergence (JSD), Invalidity Ratio (IR), LVM Score, and human Avg. Rank. Higher F1, COV, and LVM Score are preferred; lower CD, MMD, JSD, IR, and Avg. Rank are preferred.

    • CADFusion: 85.22,92.79,19.89,90.40,3.49,17.11,6.20,8.96,1.8685.22, 92.79, 19.89, 90.40, 3.49, 17.11, 6.20, 8.96, 1.86.
    • GPT-4o: 82.96,85.72,68.50,72.40,6.60,37.93,74.26,5.13,3.2282.96, 85.72, 68.50, 72.40, 6.60, 37.93, 74.26, 5.13, 3.22.
    • Text2CAD: 63.94,92.13,30.23,63.94, 92.13, 30.23, unavailable, unavailable, unavailable, 3.37,2.01,2.973.37, 2.01, 2.97; its COV, MMD, and JSD were not reported and could not be computed in the authors’ setup.

    CADFusion has the best reported value on every available metric except that Text2CAD’s extrusion F1 is close (92.13 versus 92.79). The largest visual-quality differences include LVM Score 8.96 versus 3.37 for Text2CAD and 5.13 for GPT-4o, and human Avg. Rank 1.86 versus 2.97 and 3.22, respectively. The comparison is not fully prompt-matched: Text2CAD is evaluated with its own expert-level prompts, while CADFusion and GPT-4o use the study’s prompts. The authors note that Text2CAD is sensitive to prompt detail and provide additional prompt comparisons separately.

  8. Knowl 8 — Ablations support visual feedback, alternation, and human-refined sequence captions

    empirical result

    The ablations report LVM Score (higher is better) and invalidity ratio in percent (lower is better). In the model names, SL denotes sequential learning, VF denotes visual-feedback training without a subsequent sequential-learning phase, and VFSL(nn) denotes nn alternating visual-feedback/sequential-learning rounds.

    • Sequential learning only: 7.697.69 LVM Score, 4.84%4.84\% invalidity.
    • Sequential learning on captions without human annotation: 6.566.56, 6.00%6.00\%.
    • Visual feedback without alternating sequential learning: 5.945.94, 88.87%88.87\%.
    • Visual feedback with the tested regularized preference-optimization variant: 6.216.21, 3.46%3.46\%.
    • One alternating round with human-scored preferences: 8.288.28, 17.03%17.03\%.
    • One, three, and five rounds with LVM-scored preferences: respectively 8.768.76, 4.42%4.42\%; 8.898.89, 4.21%4.21\%; and 8.968.96, 6.20%6.20\%.

    These results show that visual feedback raises the LVM score over sequential learning alone, while visual-feedback-only training has a very high invalidity ratio. More alternating rounds in the reported runs raise the LVM score, though the five-round invalidity ratio is higher than at one or three rounds. Human-scored preference data performed worse than LVM-scored data in the one-round comparison. Separately, increasing the unannotated caption dataset to about 170,000 examples did not outperform the 20,000-example human-refined dataset.

  9. Knowl 9 — Generated examples show improved instruction following across varied CAD shapes

    empirical result

    In the qualitative comparisons, CADFusion more closely follows descriptions than the tested baselines, including descriptions involving rectangles, hexagons, nested structures such as a hexagonal hole inside a cylinder, and numerical or qualitative attributes such as component counts, T-shapes, or relative length. GPT-4o often produces sequences that cannot be rendered and, when renderable, may not match the prompt. Text2CAD commonly produces well-formed regular shapes but can oversimplify complex requests or substitute basic shapes for the intended structure. CADFusion can also produce multiple variants from one prompt; the reported inference settings are temperature 0.30.3, top-pp 0.90.9, and top-kk 5050, with variation mainly in dimensions such as thickness, width, and cutout size while retaining the overall requested form.

  10. Knowl 10 — Complex item arrangements and single-view feedback limit current performance

    limitation

    The visual-feedback pipeline uses a single rendered view because the authors report that contemporary LVMs degrade when given multiple images; this restricts the visual evidence available for scoring and preference construction. CADFusion also struggles with instructions requiring substantial spatial or commonsense reasoning, including shapes that spell letters or words. The reported failure analysis identifies invalid sequences when instructions require many construction elements—for example, one case involved more than 20 loops and 50 curves—and discrepancies when prompts specify more than about eight distinct items or require complex integration of several items. In such cases, the model may miscount components, generate incorrect geometry, or fail to render.

Coverage note — The low-level, token-by-token conversion rules for every source CAD array and the supplementary gallery of additional examples are omitted because they are implementation detail or illustrations rather than separate core findings; the sequence representation, dataset construction, and principal qualitative behavior are included.

References

  1. 1.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing Systems, volume 33, pp. 1877–1901. Curran Associates, Inc., 2020.
  2. 2.Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., Nori, H., Palangi, H., Ribeiro, M. T., and Zhang, Y. Sparks of artificial general intelligence: Early experiments with gpt-4. March 2023. URL https://arxiv.org/abs/2303.12712.
  3. 3.Deng, Y., Chen, J., and Olechowski, A. What Sets Proficient and Expert Users Apart? Results of a Computer-Aided Design Experiment. Journal of Mechanical Design, 146 (1):011401, 10 2023. ISSN 1050-0472. doi: 10.1115/1.4063360. URL https://doi.org/10.1115/1.4063360.
  4. 4.Du, T., Inala, J. P., Pu, Y., Spielberg, A., Schulz, A., Rus, D., Solar-Lezama, A., and Matusik, W. Inversecsg: automatic conversion of 3d models to csg trees. ACM Trans. Graph., 37(6), December 2018. ISSN 0730-0301. doi: 10.1145/3272127.3275006. URL https://doi.org/10.1145/3272127.3275006.
  5. 5.Hong, Z., Yuan, Z., Chen, H., Zhang, Q., Huang, F., and Huang, X. Knowledge-to-SQL: Enhancing SQL generation with data expert LLM. In Ku, L.-W., Martins, A., and Srikumar, V. (eds.), Findings of the Association for Computational Linguistics: ACL 2024, pp. 10997–11008, Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.findings-acl.653. URL https://aclanthology.org/2024.findings-acl.653/.
  6. 6.Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=nZeVKeeFYf9.
  7. 7.Jayaraman, P. K., Lambourne, J. G., Desai, N., Willis, K., Sanghi, A., and Morris, N. J. W. Solidgen: An autoregressive model for direct b-rep synthesis. Transactions on Machine Learning Research, 2023. ISSN 2835-8856. URL https://openreview.net/forum?id=ZR2CDgADRo. Featured Certification.
  8. 8.Kania, K., Zieba, M., and Kajdanowicz, T. UCSG-NET-unsupervised discovering of constructive solid geometry tree. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020.
  9. 9.Kaufmann, T., Weng, P., Bengs, V., and Hüllermeier, E. A survey of reinforcement learning from human feedback, 2024. URL https://arxiv.org/abs/2312.14925.
  10. 10.Khan, M. S., Dupont, E., Ali, S. A., Cherenkova, K., Kacem, A., and Aouada, D. Cad-signet: Cad language inference from point clouds using layer-wise sketch instance guided attention. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 4713–4722, June 2024a.
  11. 11.Khan, M. S., Sinha, S., Uddin, S. T., Stricker, D., Ali, S. A., and Afzal, M. Z. Text2cad: Generating sequential CAD designs from beginner-to-expert level text prompts. In Advances in Neural Information Processing Systems, volume 37, pp. 7552–7579. Curran Associates, Inc., 2024b.
  12. 12.Lee, H., Phatale, S., Mansoor, H., Mesnard, T., Ferret, J., Lu, K., Bishop, C., Hall, E., Carbune, V., Rastogi, A., and Prakash, S. RLAIF vs. RLHF: scaling reinforcement learning from human feedback with AI feedback. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024. OpenReview.net, 2024. URL https://openreview.net/forum?id=uydQ2W41KO.
  13. 13.Li, B., Zhang, Y., Guo, D., Zhang, R., Li, F., Zhang, H., Zhang, K., Zhang, P., Li, Y., Liu, Z., et al. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326, 2024a.
  14. 14.Li, X., Song, Y., Lou, Y., and Zhou, X. CAD translator: An effective drive for text to 3d parametric computer-aided design generative modeling. In ACM Multimedia 2024, 2024b. URL https://openreview.net/forum?id=DN3722rnLd.
  15. 15.Liang, Y., He, J., Li, G., Li, P., Klimovskiy, A., Carolan, N., Sun, J., Pont-Tuset, J., Young, S., Yang, F., Ke, J., Dvijotham, K. D., Collins, K. M., Luo, Y., Li, Y., Kohlhoff, K. J., Ramachandran, D., and Navalpakkam, V. Rich human feedback for text-to-image generation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, June 16-22, 2024, pp. 19401–19411. IEEE, 2024. doi: 10.1109/CVPR52733.2024.01835. URL https://doi.org/10.1109/CVPR52733.2024.01835.
  16. 16.Liu, Y., Iter, D., Xu, Y., Wang, S., Xu, R., and Zhu, C. G-eval: NLG evaluation using gpt-4 with better human alignment. In Bouamor, H., Pino, J., and Bali, K. (eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 2511–2522, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.emnlp-main.153. URL https://aclanthology.org/2023.emnlp-main.153/.
  17. 17.Liu, Z., Lu, M., Zhang, S., Liu, B., Guo, H., Yang, Y., Blanchet, J., and Wang, Z. Provably mitigating overoptimization in rlhf: Your sft loss is implicitly an adversarial regularizer. In Globerson, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J., and Zhang, C. (eds.), Advances in Neural Information Processing Systems, volume 37, pp. 138663–138697. Curran Associates, Inc., 2024.
  18. 18.Ma, W., Chen, S., Lou, Y., Li, X., and Zhou, X. Draw step by step: Reconstructing cad construction sequences from point clouds via multimodal diffusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 27154–27163, June 2024.
  19. 19.Makatura, L., Foshey, M., Wang, B., HähnLein, F., Ma, P., Deng, B., Tjandrasuwita, M., Spielberg, A., Owens, C. E., Chen, P. Y., et al. How can large language models help humans in design and manufacturing? arXiv preprint arXiv:2307.14377, 2023.
  20. 20.Meta. The llama 3 herd of models, 2024. URL https://arxiv.org/abs/2407.21783.
  21. 21.OpenAI. Gpt-4 technical report, 2024. URL https://arxiv.org/abs/2303.08774.
  22. 22.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. Learning transferable visual models from natural language supervision. In Meila, M. and Zhang, T. (eds.), Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pp. 8748–8763. PMLR, 2021. URL http://proceedings.mlr.press/v139/radford21a.html.
  23. 23.Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C. Direct preference optimization: Your language model is secretly a reward model, 2024. URL https://arxiv.org/abs/2305.18290.
  24. 24.Seff, A., Zhou, W., Richardson, N., and Adams, R. P. Vitruvion: A generative model of parametric CAD sketches. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net, 2022. URL https://openreview.net/forum?id=Ow1C7s3UcY.
  25. 25.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G. Llama: Open and efficient foundation language models, 2023. URL https://arxiv.org/abs/2302.13971.
  26. 26.Wang, K., Zheng, J., and Zhou, Z. Neural face identification in a 2d wireframe projection of a manifold object. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 1622–1631, June 2022.
  27. 27.Willis, K. D. D., Jayaraman, P. K., Lambourne, J. G., Chu, H., and Pu, Y. Engineering sketch generation for computer-aided design. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 2105–2114, 2021. doi: 10.1109/CVPRW53098.2021.00239.
  28. 28.Wu, R., Xiao, C., and Zheng, C. Deepcad: A deep generative network for computer-aided design models. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 6772–6782, October 2021.
  29. 29.Wu, X., Huang, S., Wang, G., Xiong, J., and Wei, F. Boosting text-to-video generative model with mllms feedback. In Globerson, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J., and Zhang, C. (eds.), Advances in Neural Information Processing Systems, volume 37, pp. 139444–139469. Curran Associates, Inc., 2024.
  30. 30.Xu, X., Willis, K. D., Lambourne, J. G., Cheng, C.-Y., Jayaraman, P. K., and Furukawa, Y. Skexgen: Autoregressive generation of cad construction sequences with disentangled codebooks. In International Conference on Machine Learning, pp. 24698–24724. PMLR, 2022.
  31. 31.Xu, X., Jayaraman, P. K., Lambourne, J. G., Willis, K. D., and Furukawa, Y. Hierarchical neural coding for controllable cad model generation. In International Conference on Machine Learning, pp. 38443–38461, 2023.
  32. 32.Xu, X., Lambourne, J., Jayaraman, P., Wang, Z., Willis, K., and Furukawa, Y. Brepgen: A b-rep generative diffusion model with structured latent geometry. ACM Transactions on Graphics (TOG), 43(4):1–14, 2024.
  33. 33.Yu, F., Chen, Z., Li, M., Sanghi, A., Shayani, H., Mahdavi-Amiri, A., and Zhang, H. Capri-net: Learning compact cad shapes with adaptive primitive assembly. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 11768–11778, June 2022.
  34. 34.Yu, F., Chen, Q., Tanveer, M., Mahdavi Amiri, A., and Zhang, H. D2CSG: Unsupervised learning of compact csg trees with dual complements and dropouts. Advances in Neural Information Processing Systems, 36, 2024.
  35. 35.Zhang, J., Hou, Z., Lv, X., Cao, S., Hou, Z., Niu, Y., Hou, L., Dong, Y., Feng, L., and Li, J. Longreward: Improving long-context large language models with ai feedback, 2024. URL https://arxiv.org/abs/2410.21252.
  36. 36.Zhang, Z., Sun, S., Wang, W., Cai, D., and Bian, J. Flexcad: Unified and versatile controllable CAD generation with fine-tuned large language models. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025. URL https://openreview.net/forum?id=Z0eiiV3Yyh.
  37. 37.Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., Du, Y., Yang, C., Chen, Y., Chen, Z., Jiang, J., Ren, R., Li, Y., Tang, X., Liu, Z., Liu, P., Nie, J.-Y., and Wen, J.-R. A Survey of Large Language Models, May 2023. URL http://arxiv.org/abs/2303.18223. arXiv:2303.18223 [cs].

Citation

MLA
Wang, R., et al. “Text-to-CAD Generation Through Infusing Visual Feedback in Large Language Models”. arXiv, 2025, http://arxiv.org/abs/2501.19054v3.
APA
Wang, R., Yuan, Y., Sun, S., & Bian, J. (2025). Text-to-CAD Generation Through Infusing Visual Feedback in Large Language Models. arXiv. http://arxiv.org/abs/2501.19054v3
Chicago
Wang, R., Y. Yuan, S. Sun, and J. Bian. 2025. “Text-to-CAD Generation Through Infusing Visual Feedback in Large Language Models”. arXiv. http://arxiv.org/abs/2501.19054v3.
Harvard
Wang, R. et al. (2025) “Text-to-CAD Generation Through Infusing Visual Feedback in Large Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2501.19054v3.
Vancouver
1. Wang R, Yuan Y, Sun S, Bian J (2025) Text-to-CAD Generation Through Infusing Visual Feedback in Large Language Models. arXiv

BibTeX

@article{wang2025text,
  title = {Text-to-CAD Generation Through Infusing Visual Feedback in Large Language Models},
  author = {Wang, Ruiyu and Yuan, Yu and Sun, Shizhao and Bian, Jiang},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2501.19054v3},
  eprint = {2501.19054}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/