Benchmarking Fine-tuning and Retrieval Strategies for a Multimodal Language Model on the NRC Reactor Operator Licensing Examination
Quick Answer
The study benchmarks the 31B-parameter Gemma 4 model's performance on the NRC Reactor Operator licensing exam, revealing that supervised fine-tuning with fixed-size chunking achieved an 79.7% accuracy, surpassing all other configurations.
Quick Take
Notably, no non-fine-tuned models passed any exams, highlighting the necessity of domain-specific training for in nuclear applications.
Key Points
- Gemma 4 31B-IT model evaluated against NRC Reactor Operator licensing exams.
- Supervised fine-tuning with fixed-size chunking achieved 79.7% accuracy.
- No non-fine-tuned configurations passed any of the 14 exams.
- Preferred chunking strategy varies with model training state.
- RAFT underperformed compared to standard SFT in retrieval environments.
DeepSignal Analysis
What happened
The study benchmarks the Gemma 4 model's performance on the NRC Reactor Operator licensing exam, revealing that supervised fine-tuning with fixed-size chunking achieved 79.7% accuracy, surpassing other configurations. No non-fine-tuned models passed any exams, indicating the need for domain-specific training for LLMs in nuclear applications.
Key evidence
- The Gemma 4 model, a 31-billion-parameter multimodal model, was evaluated against the NRC Reactor Operator licensing examination.
- Supervised fine-tuning with fixed-size chunking achieved an accuracy of 79.7%, meeting the passing criterion on 8 out of 14 examinations.
- No configurations without fine-tuning passed any of the exams, underscoring the importance of domain-specific training.
Why it matters
The findings highlight the critical role of tailored training for large language models in specialized fields like nuclear power. The significant performance gap between fine-tuned and non-fine-tuned models suggests that generic models may not be sufficient for high-stakes applications. This research could inform future developments in AI applications within regulated industries, emphasizing the need for domain expertise.
Paper Resources
📖 Reader Mode
~2 min readAbstract:The integration of large language models (LLMs) into the nuclear power industry requires outputs grounded in domain-specific knowledge. This study evaluates a 31-billion-parameter open-weight multimodal model (Gemma 4 31B-IT) on its capacity to apply nuclear knowledge by benchmarking eight model-retrieval configurations against the U.S. Nuclear Regulatory Commission (NRC) Reactor Operator licensing examination. We evaluate 14 Generic Fundamentals Examinations (GFE) from the 2015-2021 March sittings (seven pressurized and seven boiling water reactor exams) using the standard 80% human passing criterion. The base model is compared against configurations utilizing supervised fine-tuning (SFT) on Gemini-distilled chain-of-thought (CoT) rationales, retrieval-augmented generation (RAG) with BM25 sparse retrieval over the U.S. Department of Energy Fundamentals Handbook, and retrieval-augmented fine-tuning (RAFT). Within the retrieval pipeline, we compare fixed-size sliding-window chunking against structure-aware chunking. The SFT configuration with fixed-size chunking RAG met the criterion on 8 of the 14 examinations, outperforming all alternatives, whereas no configuration without fine-tuning passed any. Aggregate accuracy reached 79.7%, with a confidence interval spanning the threshold, and 80.2% on PWR items specifically. Furthermore, two regularities emerged: the preferred chunking strategy reverses depending on the model's training state, and RAFT underperforms compared to standard SFT in matching search environments. These results demonstrate which combination of fine-tuning and search approaches achieves operator-level capabilities.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.22067 [cs.CL] |
| (or arXiv:2607.22067v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.22067 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: YoonPyo Lee [view email]
[v1]
Fri, 24 Jul 2026 08:10:21 UTC (90 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.