OpenRTLSet: A Fully Open-Source Dataset for Large Language Model-based Verilog Module Design
Quick Answer
OpenRTLSet is the largest open-source dataset for hardware design, featuring over 131,000 Verilog code samples.
Quick Take
It enables fine-tuning of language models like Qwen and Granite for Verilog code generation, demonstrating superior performance in hardware design tasks through open-source methodologies.
Key Points
- Dataset includes 102k Verilog modules from GitHub and 29k from translations.
- Paired natural language descriptions generated using DeepSeek-R1 model.
- Explores quantization techniques like INT4 and BF16 for performance.
- Supports fine-tuning for various language model families.
- Establishes a foundation for accessible research and commercial use.
Paper Resources
📖 Reader Mode
~2 min readAbstract:OpenRTLSet introduces the largest fully open-source dataset for hardware design, offering over 131,000 diverse Verilog code samples to the research community and industry. Our dataset uniquely combines Verilog code from GitHub repositories (102k modules), VHDL translations (5k modules), and synthesizable C/C++ translations (24k modules), all freely accessible without proprietary restrictions. Using the reasoning model DeepSeek-R1, we generated paired natural language descriptions for each code sample, enabling fine-tuning of various language model families (e.g., Qwen and Granite) for Verilog code generation. Our dataset explores multiple options, including Verilator-generated C++ files as additional context during labeling, quantization techniques (INT4 vs. BF16), and performance differences across model sizes (7B-32B parameters). OpenRTLSet demonstrates that open-source approaches can achieve superior performance in hardware design tasks, establishing a new foundation for accessible research and commercial use in this domain.
| Comments: | Accepted by ICLAD'25 |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2606.10285 [cs.CL] |
| (or arXiv:2606.10285v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2606.10285 arXiv-issued DOI via DataCite (pending registration) |
|
| Journal reference: | 2025 IEEE International Conference on LLM-Aided Design (ICLAD), Stanford, CA, USA, 2025, pp. 212-218 |
| Related DOI: | https://doi.org/10.1109/ICLAD65226.2025.00038
DOI(s) linking to related resources |
Submission history
From: Jinghua Wang [view email]
[v1]
Tue, 9 Jun 2026 01:17:46 UTC (771 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.