Text-to-CAD Generation Through Infusing Visual Feedback in Large Language Models
Ruiyu WangYu YuanShizhao SunJiang Bian
Introduces CADFusion, an alternating training framework that enables large language models to generate precise 3D CAD models from text by combining parametric sequence supervision with rendered visual feedback.
Computer-Aided Design (CAD) is essential for modern engineering and manufacturing, yet creating 3D CAD models remains a manual, time-intensive process requiring specialized technical expertise. Automated "Text-to-CAD" generation—converting natural language descriptions directly into executable parametric design sequences—offers a way to accelerate prototyping and democratize 3D modeling. However, CAD models are inherently multimodal: they exist both as sequential geometric commands and as rendered visual objects. Furthermore, because multiple distinct command sequences can produce the exact same 3D shape, training systems solely on ground-truth command sequences leads models to memorize specific sequence patterns while failing to capture global visual geometry.
The article introduces CADFusion, an artificial intelligence framework designed to evaluate and demonstrate how combining sequential code learning with automated visual feedback enables large language models to generate accurate, high-quality CAD parametric models from plain text descriptions.
The researchers utilized a fine-tuned, 8-billion-parameter open-source language model as the core engine, pairing it with a Sketch-and-Extrude representation where CAD operations are expressed as text tokens. Training alternated between two recurring stages: a sequential learning phase fine-tuning the model on 20,000 paired text-CAD examples refined by human annotators, and a visual feedback phase. To bypass the non-differentiable barrier of CAD rendering engines, the authors framed visual learning as a preference optimization task. Rendered 3D outputs were automatically scored across shape quality, component quantity, and spatial distribution using vision-language models, generating preference rankings without relying on expensive human evaluations. The framework was quantitatively evaluated against general-purpose models (GPT-4o) and specialized baseline architectures across geometric fidelity, rendering validity, visual alignment, and human expert rankings.
The experimental findings show that CADFusion substantially outperforms existing approaches across key operational metrics. First, CADFusion achieved a visual evaluation score of 8.96 out of 10 and an average human preference rank of 1.86, outperforming both general-purpose GPT-4o (score 5.13; rank 3.22) and specialized prior systems (rank 2.97). Second, the model maintained high sequence validity, rendering successfully with an invalidity ratio of only 6.20%, compared to a 74.26% failure rate for GPT-4o. Third, in geometric accuracy, CADFusion demonstrated superior point-cloud alignment with a Chamfer Distance of 19.89 and a 90.40% coverage rate, whereas prior methods struggled with complex geometries. Finally, ablation experiments confirmed that alternating between sequential training and visual feedback was essential; removing visual feedback reduced visual quality scores from 8.96 to 7.69, while visual training without alternating sequential learning caused invalidity rates to surge to 88.87% due to syntax degradation.
These results demonstrate that infusing visual preference signals into language models overcomes the core limitations of command-only text-to-CAD systems. In practice, this enables engineering organizations to rapidly generate valid, modifiable CAD assets from concise, non-expert text prompts rather than requiring tedious step-by-step drafting instructions. Because the model supports non-deterministic sampling, designers can quickly produce diverse design variations with adjustable parameters, lowering development cycle times and design iteration costs.
Organizations evaluating automated CAD workflows should consider pilot implementations using alternating sequence-and-preference optimization frameworks rather than relying on out-of-the-box language models or pure sequence generators. Development teams should prioritize curated, human-refined text prompts during initial training, as scaling unrefined synthetic datasets yielded negligible visual improvements. Future engineering initiatives must focus on incorporating multi-view visual feedback pipelines and expanding training datasets to include more intricate geometric assemblies.
These findings carry moderate-to-high confidence for standardized geometric structures and parametric primitives. However, decision-makers should exercise caution regarding current technical boundaries: the visual feedback pipeline currently relies on single-view image evaluations due to vision model constraints, and the system still struggles with highly complex spatial reasoning tasks, such as generating text-shaped geometry or coordinating more than eight distinct sub-components.
No sufficiently relevant recommendations were found.
No sufficiently relevant recommendations were found.
