ProWAFT: A ROMA-LPD Instance for Workload-Aware and Dynamic Fault Tolerance in FPGA-Based CNN Accelerators
Quick Answer
ProWAFT is a proactive fault-tolerance framework for FPGA-based CNN accelerators that uses partial reconfiguration to minimize latency, energy, and reliability risks.
Quick Take
Tested on a Xilinx Zynq UltraScale+ ZCU104 with a 500-task trace from ResNet-18 and others, it outperforms static TMR and reactive recovery, achieving high task success rates and low overhead.
Key Points
- ProWAFT employs partial reconfiguration to apply TMR selectively across reconfigurable partitions.
- Achieves lower composite cost compared to static TMR and reactive recovery methods.
- Maintains high task success rates and near-baseline throughput during fault conditions.
- Evaluated on a 500-task trace derived from ResNet-18, MobileNetV2, and EfficientNet-Lite.
- Implemented on Xilinx Zynq UltraScale+ ZCU104 platform with six reconfigurable regions.
Paper Resources
📖 Reader Mode
~2 min readAbstract:SRAM-based FPGAs provide an attractive platform for energy- and latency-constrained CNN inference at the network edge, yet transient faults can lead to silent errors that compromise reliability. Always-on redundancy (e.g., full TMR) improves correctness but incurs substantial performance and energy overhead, while reactive recovery may introduce unacceptable latency on the critical path. We propose \textbf{ProWAFT}, a proactive workload-aware fault-tolerance framework for FPGA-based CNN accelerators that uses partial reconfiguration to selectively apply TMR across reconfigurable partitions. ProWAFT quantifies workload criticality, models fault propagation and reconfiguration overhead, and selects configurations that minimize a composite objective over latency, energy, and reliability risk. Implemented on a Xilinx Zynq UltraScale+ ZCU104 platform with six reconfigurable regions and evaluated on a 500-task trace derived from ResNet-18, MobileNetV2, and EfficientNet-Lite under time-varying SEU injection, ProWAFT achieves lower composite cost than static TMR and reactive reconfiguration while maintaining high task success rate and near-baseline throughput with low online decision overhead.
| Comments: | 13 pages |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.01602 [cs.CL] |
| (or arXiv:2607.01602v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.01602 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: XinXin Chen [view email]
[v1]
Thu, 2 Jul 2026 02:04:07 UTC (10,280 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.