BetterV: Controlled Verilog Generation with Discriminative Guidance

Zehua PeiHui-Ling ZhenMingxuan YuanYu HuangBei Yu

article2024ICML182 citations

Proposes a hardware generation framework that combines domain-specific instruction tuning with discriminative guidance to produce syntactically correct Verilog code that surpasses GPT-4 on VerilogEval and optimizes downstream electronic design automation metrics.

Listen

Modern integrated circuit design faces increasing pressure from escalating complexity and the physical limits of Moore's Law. Writing hardware description languages such as Verilog is time-consuming and error-prone, which drives up production costs and slows down delivery schedules. While artificial intelligence language models offer potential for automating hardware code generation, they frequently struggle with limited training data, strict hardware correctness rules, and the multi-step optimization needs of downstream electronic design automation workflows.

The article introduces and evaluates BetterV, an automated hardware generation framework designed to produce syntactically correct, functionally valid Verilog code. The core objective is to demonstrate that domain-specific instruction tuning combined with task-specific discriminative guidance can optimize hardware implementations for subsequent engineering stages, such as logic synthesis and formal verification.

The authors curated and filtered open-source Verilog datasets, mapping code to C equivalents to transfer existing programming knowledge into language models with approximately 7 billion parameters. They enriched training through synthetic data generation and trained lightweight generative discriminators to steer generation during decoding toward desirable hardware attributes. Evaluated using industry-standard benchmarks such as VerilogEval and ANSI-C verification suites, BetterV was tested against leading commercial and open-source models, including GPT-4.

The results highlight significant performance improvements across multiple metrics. First, BetterV models achieved state-of-the-art functional correctness on the VerilogEval benchmark, outperforming GPT-4 on first-attempt pass rates by 8.1 percentage points on machine-crafted problems and 2.6 percentage points on human-crafted problems. Second, discriminative guidance reduced synthesis graph nodes by an average of 46.52% compared to reference designs, directly optimizing circuit area and complexity. Third, the framework reduced formal verification runtime by 22.45% on average against standard reference code by rewriting designs for faster mathematical satisfiability solving. Finally, guided decoding boosted syntactic compilation success rates to over 99%.

These findings demonstrate that language models can be steered beyond standard text generation into rigorous engineering optimization. In practice, early-stage optimization reduces costly design iterations between initial coding and physical synthesis, lowers engineering overhead, and speeds up time-to-market for microchips. Furthermore, achieving superior results with specialized 7-billion-parameter open-source models proves that targeted fine-tuning can rival or exceed much larger general-purpose models at a lower operational cost.

Organizations developing hardware should consider integrating task-driven discriminative guidance into their automated design flows to reduce synthesis complexity and verification bottlenecks. To implement this effectively, engineering teams must maintain verified baseline reference designs and task-specific evaluation tools to train effective discriminators. Further work should explore applying model compression methods, such as quantization and pruning, to reduce the computational overhead introduced by the discriminator during generation.

A key limitation is that BetterV requires direct access to token-level output probabilities, restricting its use to open-source language models and precluding closed-source proprietary systems. Additionally, the extra computation required by the discriminator during generation introduces latency. Nevertheless, the experimental evidence strongly supports BetterV as a robust, highly effective framework for automated hardware design optimization.

arXiv: 2402.03375
Cover for BetterV: Controlled Verilog Generation with Discriminative Guidance

Abstract

Due to the growing complexity of modern Integrated Circuits (ICs), there is a need for automated circuit design methods. Recent years have seen increasing research in hardware design language generation to facilitate the design process. In this work, we propose a Verilog generation framework, BetterV, which fine-tunes large language models (LLMs) on processed domain-specific datasets and incorporates generative discriminators for guidance on particular design demands. Verilog modules are collected, filtered, and processed from the internet to form a clean and abundant dataset. Instruct-tuning methods are specially designed to fine-tune the LLMs to understand knowledge about Verilog. Furthermore, data are augmented to enrich the training set and are also used to train a generative discriminator on particular design demands, providing guidance for the LLMs to optimize Verilog implementation. BetterV has the ability to generate syntactically and functionally correct Verilog, outperforming GPT-4 on the VerilogEval benchmark. With the help of task-specific generative discriminators, BetterV achieves remarkable improvements on various electronic design automation (EDA) downstream tasks, including netlist node reduction for synthesis and verification runtime reduction with Boolean Satisfiability (SAT) solving.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 2.1. LLMs for Verilog Generation
  • 2.2. Discriminator-guided Controllable Generation
  • 3. Algorithm
  • 3.1. Framework Overview
  • 3.2. Instruct-Tuning Data-Processing
  • 3.3. Domain-specific Instruct-Tuning
  • 3.4. Data Augmentation
  • 3.5. Generative Discriminator
  • 4. Experiments
  • 4.1. Experimental Setting
  • 4.2. Functional Correctness
  • 4.3. Customized Generation in BetterV
  • 4.3.1. SYNTHESIS NODES REDUCTION
  • 4.3.2. VERIFICATION RUNTIME REDUCTION
  • 4.4. Ablation Study
  • 4.4.1. IMPACT OF DISCRIMINATOR
  • 5. Discussion
  • 5.1. Current limitations
  • 5.2. Applicability to other domains
  • 6. Conclusion
  • Impact Statement
  • References

Knowls

  1. Knowl 1 — BetterV Framework for Controlled Verilog Generation

    model/method

    BetterV is a framework designed for controlled Hardware Description Language (HDL) generation—specifically Verilog—that combines an instruct-tuned generative large language model (LLM) with task-specific generative discriminators. The framework addresses the scarcity and complexity of domain-specific hardware data and optimizes downstream Electronic Design Automation (EDA) performance objectives directly during code generation.

    The framework consists of four primary stages:

    1. Data Processing and Alignment: Filtering open-source Verilog repositories and utilizing a Verilog-to-C translation tool (v2c) to generate paired Verilog--C programs and module definition--body pairs.

    2. Domain-Specific Instruct-Tuning: Fine-tuning pre-trained code LLMs on bi-directional Verilog↔\leftrightarrowC translation tasks (leveraging existing C knowledge to understand Verilog functional behavior) and Verilog autocompletion tasks without needing verbose, fallible natural language descriptions.

    3. Data Augmentation and Labeling: Generating diverse candidate Verilog implementations by sampling from the fine-tuned LLM at high temperature, filtering out syntax errors via EDA tools, and labeling implementations as desired (cc) or undesired (cˉ\bar{c}) according to downstream metrics (e.g., syntax correctness, equivalence, synthesis netlist node count, or formal verification runtime).

    4. Discriminator Guidance: Training a lightweight class-conditional generative discriminator using a joint generative-discriminative loss, which subsequently guides the generative LLM during decoding via Bayes' rule.

  2. Knowl 2 — Domain-Specific Verilog Instruct-Tuning via C Code Knowledge Transfer and Autocompletion

    model/method

    To overcome the global scarcity of high-quality Verilog corpora and the fallibility of machine-generated natural language code descriptions, BetterV employs a dual-task domain-specific instruct-tuning strategy on open-source Verilog data:

    1. Verilog--C Cross-Translation Alignment: Verilog modules are translated to C using a Verilog-to-C translation tool (v2c). The generative LLM is trained on paired instructions to translate Verilog into C and C into Verilog. Because LLMs already possess extensive pre-trained knowledge of general-purpose programming languages like C, the translated C code serves as a clean, structured functional description that grounds the semantics of the hardware description.

    2. Verilog Module Autocompletion: The LLM is provided with the module definition (including module header, parameter lists, and input/output port declarations) in the instruction prompt and is trained to complete the full module implementation (module body).

    The training objective optimizes the standard auto-regressive negative log-likelihood over token sequences x={x1,…,xn}x = \{x_1, \dots, x_n\} across the training dataset D\mathcal{D}:

    L=−1∣D∣∑i=1∣D∣1ni∑t=1nilog⁡pθ(xti∣x<ti)\mathcal{L} = -\frac{1}{|\mathcal{D}|} \sum_{i=1}^{|\mathcal{D}|} \frac{1}{n_i} \sum_{t=1}^{n_i} \log p_\theta(x_t^i \mid x_{<t}^i)

    where θ\theta denotes the model parameters, xtix_t^i is the tt-th token of sequence ii, and nin_i is the sequence length.

  3. Knowl 3 — Hybrid Loss Formulation for Training the Generative Discriminator

    equation

    In the BetterV framework, the task-specific generative discriminator is modeled as a class-conditional language model (CC-LM) conditioned on a binary control code c′∈{c,cˉ}c' \in \{c, \bar{c}\}, where cc represents the desired attribute and cˉ\bar{c} represents the undesired attribute. The discriminator is trained using a hybrid objective combining a generative loss Lg\mathcal{L}_g and a discriminative loss Ld\mathcal{L}_d.

    The generative loss Lg\mathcal{L}_g trains the model to predict the next token given the conditioning code c′c' across sequences in dataset D\mathcal{D}:

    Lg=−1∣D∣∑i=1∣D∣1n∑t=1nlog⁡pθ(xti∣x<ti,c′)\mathcal{L}_g = -\frac{1}{|\mathcal{D}|} \sum_{i=1}^{|\mathcal{D}|} \frac{1}{n} \sum_{t=1}^n \log p_\theta(x_t^i \mid x_{<t}^i, c')

    The posterior probability that a sequence prefix x1:tx_{1:t} belongs to class label cy∈{c,cˉ}c_y \in \{c, \bar{c}\} is computed via Bayes' rule with length-normalized likelihoods:

    pθ(cy∣x1:t)=p(cy) pθ(x1:t∣cy)α/t∑c′∈{c,cˉ}p(c′) pθ(x1:t∣c′)α/tp_\theta(c_y \mid x_{1:t}) = \frac{p(c_y) \, p_\theta(x_{1:t} \mid c_y)^{\alpha / t}}{\sum_{c' \in \{c, \bar{c}\}} p(c') \, p_\theta(x_{1:t} \mid c')^{\alpha / t}}

    where α\alpha is a learnable scale parameter, p(c)=ebc∑c′ebc′p(c) = \frac{e^{b_c}}{\sum_{c'} e^{b_{c'}}} with learnable class bias bcb_c, and tt is the current sequence length.

    The discriminative loss Ld\mathcal{L}_d ensures the model accurately classifies full sequences x1:nix^i_{1:n} into their ground-truth labels cyic_y^i:

    Ld=−1∣D∣∑i=1∣D∣log⁡pθ(cyi∣x1:ni)\mathcal{L}_d = -\frac{1}{|\mathcal{D}|} \sum_{i=1}^{|\mathcal{D}|} \log p_\theta(c_y^i \mid x_{1:n}^i)

    The total training loss Ltotal\mathcal{L}_{\text{total}} is a convex combination weighted by hyperparameter λ∈[0,1]\lambda \in [0, 1]:

    Ltotal=λLg+(1−λ)Ld\mathcal{L}_{\text{total}} = \lambda \mathcal{L}_g + (1 - \lambda)\mathcal{L}_d
  4. Knowl 4 — Discriminator-Guided Weighted Decoding and Token Filtering Algorithm

    algorithm

    BetterV guides next-token generation of an LLM using the trained generative discriminator through weighted decoding combined with dual-criterion token filtering.

    Input: Generative LLM pLLMp_{\text{LLM}}, Generative Discriminator pθp_\theta, Desired control code cc, Prompt prefix x<tx_{<t}, Vocabulary set V\mathcal{V}, Influence weight ww, Cumulative probability threshold ρ\rho, Probability cutoff threshold τ\tau
    Output: Sampled next token xtx_t
    for each token v∈Vv \in \mathcal{V} do
        Calculate candidate sequence x1:t=(x<t,v)x_{1:t} = (x_{<t}, v)
        Compute discriminator posterior pθ(c∣v,x<t)p_\theta(c \mid v, x_{<t}) using Bayes' rule
        Compute weighted probability pw(v∣x<t,c)∝pLLM(v∣x<t)⋅[pθ(c∣v,x<t)]wp_w(v \mid x_{<t}, c) \propto p_{\text{LLM}}(v \mid x_{<t}) \cdot \left[p_\theta(c \mid v, x_{<t})\right]^w
    end for
    Rank all tokens v∈Vv \in \mathcal{V} in descending order of pθ(c∣v,x<t)p_\theta(c \mid v, x_{<t}) to obtain ordered list Vrank\mathcal{V}_{\text{rank}}
    Find minimum index mm such that ∑j=1mpw(Vrank[j]∣x<t,c)≥ρ\sum_{j=1}^m p_w(\mathcal{V}_{\text{rank}}[j] \mid x_{<t}, c) \ge \rho
    Define cumulative mass set Vm=Vrank[1:m]\mathcal{V}_m = \mathcal{V}_{\text{rank}}[1:m]
    Define threshold set Vτ={v∈V∣pθ(c∣v,x<t)>τ}\mathcal{V}_\tau = \{v \in \mathcal{V} \mid p_\theta(c \mid v, x_{<t}) > \tau\}
    Construct candidate filtering set Vk=Vτ∪Vm\mathcal{V}_k = \mathcal{V}_\tau \cup \mathcal{V}_m
    Normalize pw(xt∣x<t,c)p_w(x_t \mid x_{<t}, c) over xt∈Vkx_t \in \mathcal{V}_k
    Sample xt∼pw(xt∣x<t,c)x_t \sim p_w(x_t \mid x_{<t}, c) restricted to Vk\mathcal{V}_k
    return xtx_t

    The hyperparameter ww controls the strength of discriminator guidance. Candidate filtering retains high-quality tokens under the discriminator without excluding tokens that meet the probability threshold τ\tau or the cumulative mass threshold ρ\rho.

  5. Knowl 5 — Verilog Functional Correctness Performance on VerilogEval Benchmark

    data/table

    The functional correctness of Verilog code generated by BetterV variants was evaluated against existing open-source and proprietary models on the VerilogEval benchmark, comprising machine-generated and human-crafted hardware problems. Evaluations used the unbiased estimator for pass@k\text{pass}@k with k∈{1,5,10}k \in \{1, 5, 10\} based on n=20n=20 samples per problem.

    Model VerilogEval-machine VerilogEval-human
    pass@1 pass@5 pass@10 pass@1 pass@5 pass@10
    GPT-3.5 46.7 69.1 74.1 26.7 45.8 51.7
    GPT-4 60.0 70.6 73.5 43.5 55.8 58.9
    CodeLlama (7B) 43.1 47.1 47.7 18.2 22.7 24.3
    DeepSeek (6.7B) 52.2 55.4 56.8 30.2 33.9 34.9
    CodeQwen (7B) 46.5 54.9 56.4 22.5 26.1 28.0
    ChipNeMo 43.4 - - 22.4 - -
    Thakur et al. 44.0 52.6 59.2 30.3 43.9 49.6
    VerilogEval 46.2 67.3 73.7 28.8 45.9 52.3
    RTLCoder-Mistral 62.5 72.2 76.6 36.7 45.5 49.2
    RTLCoder-DeepSeek 61.2 76.5 81.8 41.6 50.1 53.4
    BetterV-CodeLlama 64.2 75.4 79.1 40.9 50.0 53.3
    BetterV-DeepSeek 67.8 79.1 84.0 45.9 53.3 57.6
    BetterV-CodeQwen 68.1 79.4 84.5 46.1 53.7 58.2

    BetterV-CodeQwen achieves state-of-the-art results across both benchmarks, outperforming GPT-4 on pass@1\text{pass}@1 by 8.1 percentage points on VerilogEval-machine (68.1% vs. 60.0%) and by 2.6 percentage points on VerilogEval-human (46.1% vs. 43.5%), without prompt-engineering strategies.

  6. Knowl 6 — Synthesis Netlist Node Reduction Guided by Generative Discriminator

    data/table

    BetterV incorporates a downstream task-specific generative discriminator to guide the LLM to rewrite Verilog modules into implementations that yield fewer And-Inverter-Graph (AIG) netlist nodes after logic synthesis. Netlist node counts were measured using Yosys synthesis commands (proc; aigmap; stat) on problems from the VerilogEval-human dataset.

    Problem Ref BetterV-base BetterV Com Base (%) Com Ref (%)
    ece241_2013_q8 657 333.5 255.3 23.44% 61.14%
    m2041_q6 1370 692.7 685.6 1.03% 49.95%
    counter_2bc 673 666.2 518.9 22.11% 22.89%
    review2015_count1k 487 493.4 402.6 18.44% 17.33%
    timer 498 294.3 247.3 15.97% 50.34%
    edgedetect2 58 189.9 47.4 75.03% 18.27%
    counter1to10 325 266.3 240.3 9.76% 26.06%
    2013_q2afsm 826 308.8 296.6 3.95% 64.09%
    dff8p 50 42.3 37.8 10.63% 24.40%
    fsm3comb 844 167.9 104.4 37.82% 87.63%
    rule90 6651 12435.6 4536.9 63.52% 31.79%
    mux256to1v 2376 2439.6 557.2 77.16% 76.54%
    fsm2 389 186.53 121.9 34.65% 68.66%
    fsm2s 396 163.7 144.1 11.97% 63.61%
    ece241_2013_q4 2222 1789.5 897.4 49.85% 59.61%
    conwaylife 43794 547400.3 27037.4 95.06% 38.26%
    count_clock 3187 2497.5 2222.2 11.02% 30.27%
    countbcd 1589 932.0 849.3 8.87% 46.55%

    The reported node counts are averages over generated completions. "Com Base" and "Com Ref" denote the percentage reduction in AIG nodes achieved by discriminator-guided BetterV relative to BetterV-base (the instruct-tuned model without discriminator guidance) and the reference Verilog implementation, respectively. Across the benchmark suite, BetterV achieves an average node count reduction of 46.52% compared to the reference designs and 31.68% compared to BetterV-base.

  7. Knowl 7 — Boolean Satisfiability Formal Verification Runtime Reduction

    data/table

    BetterV applies discriminator guidance to optimize Verilog/SystemVerilog designs for formal verification efficiency, aiming to reduce the Boolean Satisfiability (SAT) solving runtime required to verify assertion safety. Using designs from the ANSI-C benchmark, SAT solving times were measured using Yosys (hierarchy; proc; opt; sat -verify -seq 100 -tempinduct -prove-asserts) to prove all embedded assertions up to 100 time steps.

    Design Ref (s) BetterV-base (s) BetterV (s) Com Base (%) Com Ref (%)
    b03 1.233 1.252 0.857 31.54% 30.49%
    b06 0.099 0.083 0.078 6.02% 21.21%
    Spinner 1.577 1.343 1.064 20.77% 32.53%
    traffic_light_example 0.583 0.497 0.480 3.42% 17.67%
    Rotate 1.153 1.126 1.034 8.17% 10.32%

    "Com Base" and "Com Ref" represent the percentage reduction in SAT verification runtime achieved by BetterV (guided by the verification runtime discriminator) compared to BetterV-base (without discriminator guidance) and the reference implementation, respectively. On average, BetterV reduced SAT verification time by 22.45% compared to the reference implementations and by 13.99% compared to BetterV-base.

  8. Knowl 8 — Ablation Study on Discriminator Guidance for Syntactic and Functional Correctness

    empirical result

    Ablation experiments evaluate the impact of plugging in a task-specific generative discriminator during generation on the VerilogEval-human dataset for both functional correctness (evaluated using Yosys eqy for logical equivalence) and syntactic correctness (evaluated using Yosys generic synthesis script prep):

    1. Functional Correctness Impact:

      • For the raw pre-trained CodeLlama-7B-Instruct, adding the equivalence discriminator (CodeLlama + Dis) increases pass@1\text{pass}@1 from 18.2% to 20.3%, pass@5\text{pass}@5 from 22.7% to 24.1%, and pass@10\text{pass}@10 from 24.3% to 24.7%.
      • For BetterV-CodeLlama-base (instruct-tuned without discriminator), adding the discriminator (BetterV-CodeLlama) increases pass@1\text{pass}@1 from 40.0% to 40.9%, pass@5\text{pass}@5 from 49.5% to 50.0%, and pass@10\text{pass}@10 from 53.0% to 53.3%.
    2. Syntactic Correctness Impact:

      • For CodeLlama-7B-Instruct, adding the syntax discriminator increases syntactic pass@1\text{pass}@1 from 41.5% to 49.1% (+7.6 percentage points), pass@5\text{pass}@5 from 50.8% to 58.4%, and pass@10\text{pass}@10 from 53.8% to 61.1%.
      • For BetterV-CodeLlama-base, adding the syntax discriminator improves syntactic pass@1\text{pass}@1 from 82.6% to 87.1% (+4.5 percentage points), pass@5\text{pass}@5 from 97.6% to 98.2%, and pass@10\text{pass}@10 from 99.2% to 99.3%.

    These results show that the trained discriminator can guide both base pre-trained models and instruct-tuned models as a plug-and-play module, provided they share the same vocabulary.

  9. Knowl 9 — Limitations of the BetterV Discriminative Guidance Framework

    limitation

    The BetterV framework has two primary limitations:

    1. Inapplicability to Black-Box Closed-Source Models: The discriminator-guided decoding mechanism relies on computing conditional token-level probabilities and modifying the generative model's next-token distribution pLLM(xt∣x<t)p_{\text{LLM}}(x_t \mid x_{<t}) via Bayes' rule. Consequently, it cannot be directly applied to closed-source API-only LLMs (such as GPT-4) where token-level logits or sampling distributions are inaccessible.

    2. Inference-Time Computational Overhead: Incorporating a generative discriminator requires evaluating next-token candidate probabilities under both control codes at each auto-regressive decoding step. Although the discriminator uses a smaller architecture (e.g., 1.1B–1.3B parameters), it introduces non-negligible inference overhead that necessitates model compression techniques such as pruning or quantization for practical deployment.

Coverage note — No substantial contributed material was omitted.

References

  1. 1.Bai, J., Bai, S., Chu, Y., Cui, Z., Dang, K., Deng, X., Fan, Y., Ge, W., Han, Y., Huang, F., et al. Qwen technical report. arXiv preprint arXiv:2309.16609, 2023.
  2. 2.Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021.
  3. 3.Dathathri, S., Madotto, A., Lan, J., Hung, J., Frank, E., Molino, P., Yosinski, J., and Liu, R. Plug and Play Language Models: A Simple Approach to Controlled Text Generation. In International Conference on Learning Representations (ICLR), 2020.
  4. 4.Dehaerne, E., Dey, B., Halder, S., and De Gendt, S. A Deep Learning Framework for Verilog Autocompletion Towards Design and Verification Automation. arXiv preprint arXiv:2304.13840, 2023.
  5. 5.Guo, D., Zhu, Q., Yang, D., Xie, Z., Dong, K., Zhang, W., Chen, G., Bi, X., Wu, Y., Li, Y., et al. Deepseek-coder: When the large language model meets programming–the rise of code intelligence. arXiv preprint arXiv:2401.14196, 2024.
  6. 6.He, Z., Wu, H., Zhang, X., Yao, X., Zheng, S., Zheng, H., and Yu, B. ChatEDA: A large language model powered autonomous agent for EDA. In ACM/IEEE Workshop on Machine Learning CAD (MLCAD), pp. 1–6. IEEE, 2023.
  7. 7.Holtzman, A., Buys, J., Forbes, M., Bosselut, A., Golub, D., and Choi, Y. Learning to write with cooperative discriminators. arXiv preprint arXiv:1805.06087, 2018.
  8. 8.Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021.
  9. 9.Keskar, N. S., McCann, B., Varshney, L. R., Xiong, C., and Socher, R. Ctrl: A conditional transformer language model for controllable generation. arXiv preprint arXiv:1909.05858, 2019.
  10. 10.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  11. 11.Krause, B., Gotmare, A. D., McCann, B., Keskar, N. S., Joty, S., Socher, R., and Rajani, N. F. Gedi: Generative discriminator guided sequence generation. arXiv preprint arXiv:2009.06367, 2020.
  12. 12.Liu, A., Sap, M., Lu, X., Swayamdipta, S., Bhagavatula, C., Smith, N. A., and Choi, Y. DExperts: Decoding-time controlled text generation with experts and anti-experts. arXiv preprint arXiv:2105.03023, 2021.
  13. 13.Liu, M., Ene, T.-D., Kirby, R., Cheng, C., Pinckney, N., Liang, R., Alben, J., Anand, H., Banerjee, S., Bayraktaroglu, I., et al. ChipNeMo: Domain-Adapted LLMs for Chip Design. arXiv preprint arXiv:2311.00176, 2023a.
  14. 14.Liu, M., Pinckney, N., Khailany, B., and Ren, H. VerilogEval: Evaluating Large Language Models for Verilog Code Generation. In IEEE/ACM International Conference on Computer-Aided Design (ICCAD), pp. 1–8. IEEE, 2023b.
  15. 15.Liu, S., Fang, W., Lu, Y., Zhang, Q., Zhang, H., and Xie, Z. RTLCoder: Outperforming GPT-3.5 in Design RTL Generation with Our Open-Source Dataset and Lightweight Solution. arXiv preprint arXiv:2312.08617, 2023c.
  16. 16.Loshchilov, I. and Hutter, F. Sgdr: Stochastic gradient descent with warm restarts. arXiv preprint arXiv:1608.03983, 2016.
  17. 17.Lu, Y., Liu, S., Zhang, Q., and Xie, Z. RTLLM: An open-source benchmark for design rtl generation with large language model. arXiv preprint arXiv:2308.05345, 2023.
  18. 18.Mukherjee, R., Kroening, D., and Melham, T. Hardware verification using software analyzers. In IEEE Computer Society Annual Symposium on VLSI, pp. 7–12. IEEE, 2015. ISBN 978-1-4799-8719-1.
  19. 19.Mukherjee, R., Tautschnig, M., and Kroening, D. v2c – a Verilog to C translator tool. In Tools and Algorithms for the Construction and Analysis of Systems (TACAS), volume 9636 of LNCS, pp. 580–586. Springer, 2016. ISBN 978-3-662-49673-2.
  20. 20.Nijkamp, E., Hayashi, H., Xiong, C., Savarese, S., and Zhou, Y. Codegen2: Lessons for training llms on programming and natural languages. arXiv preprint arXiv:2305.02309, 2023.
  21. 21.Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y. Zero: Memory optimizations toward training trillion parameter models. In ACM/IEEE Supercomputing Conference (SC), pp. 1–16. IEEE, 2020.
  22. 22.Roziere, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X. E., Adi, Y., Liu, J., Remez, T., Rapin, J., et al. Code llama: Open foundation models for code. arXiv preprint arXiv:2308.12950, 2023.
  23. 23.Scialom, T., Dray, P.-A., Lamprier, S., Piwowarski, B., and Staiano, J. Discriminative adversarial search for abstractive summarization. In International Conference on Machine Learning (ICML), pp. 8555–8564. PMLR, 2020.
  24. 24.Thakur, S., Ahmad, B., Fan, Z., Pearce, H., Tan, B., Karri, R., Dolan-Gavitt, B., and Garg, S. Benchmarking Large Language Models for Automated Verilog RTL Code Generation. In IEEE/ACM Proceedings Design, Automation and Test in Eurpoe (DATE), pp. 1–6. IEEE, 2023.
  25. 25.Wolf, C. Yosys open synthesis suite. https://yosyshq.net/yosys/.
  26. 26.Yao, Z., Aminabadi, R. Y., Ruwase, O., Rajbhandari, S., Wu, X., Awan, A. A., Rasley, J., Zhang, M., Li, C., Holmes, C., Zhou, Z., Wyatt, M., Smith, M., Kurilenko, L., Qin, H., Tanaka, M., Che, S., Song, S. L., and He, Y. DeepSpeed-Chat: Easy, Fast and Affordable RLHF Training of ChatGPT-like Models at All Scales. arXiv preprint arXiv:2308.01320, 2023.
  27. 27.Zhang, P., Zeng, G., Wang, T., and Lu, W. TinyLlama: An Open-Source Small Language Model, 2024.

Citation

MLA
Pei, Z., et al. “BetterV: Controlled Verilog Generation with Discriminative Guidance”. arXiv, 2024, http://arxiv.org/abs/2402.03375v3.
APA
Pei, Z., Zhen, H.-L., Yuan, M., Huang, Y., & Yu, B. (2024). BetterV: Controlled Verilog Generation with Discriminative Guidance. arXiv. http://arxiv.org/abs/2402.03375v3
Chicago
Pei, Z., H.-L. Zhen, M. Yuan, Y. Huang, and B. Yu. 2024. “BetterV: Controlled Verilog Generation with Discriminative Guidance”. arXiv. http://arxiv.org/abs/2402.03375v3.
Harvard
Pei, Z. et al. (2024) “BetterV: Controlled Verilog Generation with Discriminative Guidance”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.03375v3.
Vancouver
1. Pei Z, Zhen H-L, Yuan M, Huang Y, Yu B (2024) BetterV: Controlled Verilog Generation with Discriminative Guidance. arXiv

BibTeX

@article{pei2024betterv,
  title = {BetterV: Controlled Verilog Generation with Discriminative Guidance},
  author = {Pei, Zehua and Zhen, Hui-Ling and Yuan, Mingxuan and Huang, Yu and Yu, Bei},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.03375v3},
  eprint = {2402.03375}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/