BetterV: Controlled Verilog Generation with Discriminative Guidance
Zehua PeiHui-Ling ZhenMingxuan YuanYu HuangBei Yu
Proposes a hardware generation framework that combines domain-specific instruction tuning with discriminative guidance to produce syntactically correct Verilog code that surpasses GPT-4 on VerilogEval and optimizes downstream electronic design automation metrics.
Modern integrated circuit design faces increasing pressure from escalating complexity and the physical limits of Moore's Law. Writing hardware description languages such as Verilog is time-consuming and error-prone, which drives up production costs and slows down delivery schedules. While artificial intelligence language models offer potential for automating hardware code generation, they frequently struggle with limited training data, strict hardware correctness rules, and the multi-step optimization needs of downstream electronic design automation workflows.
The article introduces and evaluates BetterV, an automated hardware generation framework designed to produce syntactically correct, functionally valid Verilog code. The core objective is to demonstrate that domain-specific instruction tuning combined with task-specific discriminative guidance can optimize hardware implementations for subsequent engineering stages, such as logic synthesis and formal verification.
The authors curated and filtered open-source Verilog datasets, mapping code to C equivalents to transfer existing programming knowledge into language models with approximately 7 billion parameters. They enriched training through synthetic data generation and trained lightweight generative discriminators to steer generation during decoding toward desirable hardware attributes. Evaluated using industry-standard benchmarks such as VerilogEval and ANSI-C verification suites, BetterV was tested against leading commercial and open-source models, including GPT-4.
The results highlight significant performance improvements across multiple metrics. First, BetterV models achieved state-of-the-art functional correctness on the VerilogEval benchmark, outperforming GPT-4 on first-attempt pass rates by 8.1 percentage points on machine-crafted problems and 2.6 percentage points on human-crafted problems. Second, discriminative guidance reduced synthesis graph nodes by an average of 46.52% compared to reference designs, directly optimizing circuit area and complexity. Third, the framework reduced formal verification runtime by 22.45% on average against standard reference code by rewriting designs for faster mathematical satisfiability solving. Finally, guided decoding boosted syntactic compilation success rates to over 99%.
These findings demonstrate that language models can be steered beyond standard text generation into rigorous engineering optimization. In practice, early-stage optimization reduces costly design iterations between initial coding and physical synthesis, lowers engineering overhead, and speeds up time-to-market for microchips. Furthermore, achieving superior results with specialized 7-billion-parameter open-source models proves that targeted fine-tuning can rival or exceed much larger general-purpose models at a lower operational cost.
Organizations developing hardware should consider integrating task-driven discriminative guidance into their automated design flows to reduce synthesis complexity and verification bottlenecks. To implement this effectively, engineering teams must maintain verified baseline reference designs and task-specific evaluation tools to train effective discriminators. Further work should explore applying model compression methods, such as quantization and pruning, to reduce the computational overhead introduced by the discriminator during generation.
A key limitation is that BetterV requires direct access to token-level output probabilities, restricting its use to open-source language models and precluding closed-source proprietary systems. Additionally, the extra computation required by the discriminator during generation introduces latency. Nevertheless, the experimental evidence strongly supports BetterV as a robust, highly effective framework for automated hardware design optimization.
- Paper: Evaluating Large Language Models Trained on Code, Mark Chen et al. (2021). This foundational paper establishes the standard methodologies for evaluating large language models on functional code generation via unit-test execution.
- Paper: CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis, Erik Nijkamp et al. (2022). This work introduces open-source autoregressive code generation models and multi-turn synthesis architectures that underpin open code-generation research.
- Paper: Magicoder: Empowering Code Generation with OSS-Instruct, Yuxiang Wei et al. (2024). This paper demonstrates synthetic instruction-tuning techniques and compact open-source model optimization for code generation, foundational to BetterV's training approach.
- Paper: Textbooks Are All You Need, Suriya Gunasekar et al. (2023). This work establishes how high-quality synthetic data enables compact language models to achieve state-of-the-art coding performance with minimal parameter scale.
- Paper: Natural Language to Code Translation with Execution, Freda Shi et al. (2022). This study introduces execution-guided decoding and consensus selection at inference time, providing foundational concepts for guided program synthesis.
- Paper: Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation, Jiawei Liu et al. (2023). This paper establishes the importance of rigorous functional verification and adversarial test-case generation to reliably assess generated code correctness.
- Paper: AnalogCoder: Analog Circuit Design via Training-Free Code Generation, Yao Lai et al. (2025). This work extends automated hardware code generation from digital Verilog synthesis into the domain of analog circuit design using agentic flows.
- Paper: Generative Verifiers: Reward Modeling as Next-Token Prediction, Lunjun Zhang et al. (2025). This text explores generative reward modeling and verification strategies that advance beyond token-level discriminative guidance during reasoning and synthesis.
- Paper: Process Reward Model with Q-value Rankings, Wendi Li et al. (2025). This research continues the development of intermediate step verification by formulating process reward modeling as sequential value estimation.
- Paper: The Verification Horizon: No Silver Bullet for Coding Agent Rewards, Binghai Wang et al. (2026). This work examines the broader limits, failure modes, and reward hacking challenges of verifier-guided coding agents across downstream engineering workflows.
- Paper: VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models, Sen Xu et al. (2026). This book analyzes how small language models can be pushed to frontier levels of verifiable reasoning and algorithmic code generation.
