
NVIDIA Nemotron 3 Ultra Leads Open Models on Accuracy and Efficiency in Agentic RTL Coding
Quick Answer
NVIDIA's Nemotron 3 Ultra, in conjunction with the ACE-RTL agent, achieves a 100% pass rate on RTL tasks, outperforming competitors like GLM 5.2 and Kimi K2.6 while using 28% fewer tokens.
Quick Take
This efficiency is crucial for modern chip design, where iterative feedback is essential for accuracy.
Key Points
- Nemotron 3 Ultra achieves a 97.1% average pass rate across nine CVDP task categories.
- The model uses 6,629 tokens per iteration, significantly lower than competitors.
- ACE-RTL agent enhances debugging efficiency, improving pass rates across all models evaluated.
- Nemotron 3 Ultra is designed for long-context reasoning, crucial for RTL workflows.
- The hybrid architecture allows for up to 5X higher throughput and 30% lower costs.
DeepSignal Analysis
What happened
NVIDIA's Nemotron 3 Ultra, paired with the ACE-RTL agent, achieved a 100% pass rate on RTL tasks, surpassing competitors like GLM 5.2 and Kimi K2.6. It also utilized 28% fewer tokens than GLM 5.2, demonstrating significant efficiency in RTL coding. This performance is particularly relevant in modern chip design, where iterative feedback is essential for accuracy.
Key evidence
- Nemotron 3 Ultra achieved a 100% pass rate on RTL tasks, while GLM 5.2 and Kimi K2.6 had lower rates of 94.3% and 97.1%, respectively.
- The model uses an average of 6,629 tokens per iteration, which is 28% fewer than GLM 5.2's 9,156 tokens and 71% fewer than Kimi K2.6's 22,579 tokens.
- The CVDP benchmark evaluates LLMs on realistic RTL tasks, reflecting practical hardware design problems more accurately than previous benchmarks.
Why it matters
The ability of Nemotron 3 Ultra to achieve high accuracy with lower token usage is significant for chip design, where efficiency and speed are critical. The iterative nature of RTL coding means that reducing the number of tokens used per iteration can lead to faster results and more attempts within the same computational budget. This can enhance productivity for engineers working on complex hardware designs, making the tool more practical for real-world applications.
Source Excerpt
Modern chip design is increasingly limited by engineering time. Register transfer level (RTL) development and verification require specialized hardware…
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from NVIDIA Developer Blog
See more →
Synthetic Data Generation for Financial AI Research with NVIDIA NeMo
NVIDIA's NeMo pipeline generates 502,536 unique financial news headlines in 82 iterations, addressing data imbalance in financial NLP. The iterative approach uses semantic deduplication and category-weighted sampling to enhance diversity and relevance in generated content.

