Data Quality over Capacity: Internalizing Documents into LoRA Adapters for Closed-Book QA
Quick Answer
This study demonstrates that the quality of training data significantly impacts closed-book QA accuracy when documents are internalized into a 4-bit Gemma-4-e4b model via LoRA.
Quick Take
A single curation pass improved accuracy from 57.7% to 85.7% on a 15-document corpus, outperforming traditional retrieval methods like BM25-.
Key Points
- Training data quality is the key factor for closed-book QA accuracy.
- A single curation pass improved accuracy from 57.7% to 85.7%.
- Internalized adapter achieved 84.2% recall, outperforming BM25-RAG.
- Capacity must grow with corpus size for optimal performance.
- Study includes three misdiagnoses as a case study in training.
Paper Resources
📖 Reader Mode
~2 min readAbstract:We study baking documents directly into the weights of a 4-bit Gemma-4-e4b model via LoRA, so a system can answer questions about a corpus closed-book: no retrieval and no context-window budget. Across roughly 100 training runs from single documents to a 99-document corpus, we find that once adapter capacity is adequate, training-data quality is the dominant lever on closed-book accuracy, outweighing LoRA rank, learning rate, and two alternative architectures combined; capacity itself is a hard gate below which no data intervention helps. A single curation pass (shortening gold answers to canonical 1-6 word spans and dropping trivia) moved closed-book accuracy from 57.7% to 85.7% on a 15-document corpus, a larger jump than any architectural change. We confirm a capacity trend (rank must grow with corpus size) entangled with a coupling between rank and learning rate that we initially misdiagnosed. On a 15-document slice we add a real retrieval baseline: the internalized adapter (84.2% recall) beats a BM25-RAG pipeline with a base reader (58.9%) and even a realistic gold-chunk oracle (65.6%) at lower latency. We report the full arc, including three misdiagnoses, as a case study in debugging LLM training empirically.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.21861 [cs.CL] |
| (or arXiv:2607.21861v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.21861 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Joan Figuerola Hurtado [view email]
[v1]
Thu, 23 Jul 2026 23:14:59 UTC (12 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.