A Classifier That Teaches Itself: Self-Improving, Frozen-gate Training (SIFT) for Dynamic Document Classification
Quick Answer
This paper shows that SIFT (Self-Improving, Frozen-gate Training) introduces a dynamic document classification service that improves over time by utilizing a CPU-bound pipeline and a LightGBM head, escalating only low-confidence cases to an LLM judge.
Quick Take
This approach reduces labeling costs and enhances accuracy while ensuring safety through a two-part promotion mechanism, allowing for autonomous retraining without human intervention.
Key Points
- SIFT uses a SPLADE sparse encoder and LightGBM for efficient classification.
- Low-confidence pages are escalated to an judge for accurate labeling.
- The model learns from production traffic, reducing upfront annotation efforts.
- Safety is ensured with a two-part promote gate to prevent regression.
- Marginal labeling costs trend towards zero with continuous self-improvement.
DeepSignal Analysis
What happened
The paper presents SIFT, a dynamic document classification service that utilizes a CPU-bound pipeline and a LightGBM head. It escalates low-confidence cases to a large language model (LLM) for judgment, allowing for continuous improvement without extensive human labeling efforts.
Key evidence
- SIFT employs a SPLADE sparse encoder feeding into a LightGBM head, which is designed to be cost-effective and CPU-bound.
- The system escalates only low-confidence cases to an LLM judge, which helps in continuously updating the labeled corpus with minimal human intervention.
- A two-part promotion mechanism is implemented to ensure safety, involving a critical-label F1 regression check and a frozen golden regression set.
Why it matters
SIFT addresses the challenges of document classification in enterprise settings, where traditional labeling processes can be resource-intensive. By automating the retraining process and reducing labeling costs, it has the potential to enhance classification accuracy over time. This could lead to more efficient document management systems in various industries.
What to watch
Paper Resources
📖 Reader Mode
~2 min readAbstract:Document classification is a solved problem in the laboratory and an unsolved one in the enterprise. The blocker is rarely model architecture; it is the labeling project that must precede a model and the institutional fear of letting a model retrain itself once one exists. We present SIFT (Self-Improving, Frozen-gate Training), a dynamic classifier service, which attacks both. SIFT serves classification from a deliberately cheap, CPU-bound pipeline, a SPLADE sparse encoder feeding a LightGBM head, and escalates only the low-confidence minority of pages to an LLM judge. The judge's verdicts are written back into a labeled corpus, so the expensive model continuously teaches the cheap one: the escalation rate falls, the corpus grows from production traffic rather than from an up-front annotation effort, and accuracy compounds with use. Onboarding a new document family requires only a declarative bundle, label space, anchor phrases, and a judge glossary, not a labeling project. The harder problem is safety: an autonomously retraining classifier can silently regress. SIFT resolves this with a two-part promote gate, a critical-label F1 regression check plus a frozen golden regression set the model is never trained on, either of which vetoes promotion. This turns "retrain monthly without a human" from reckless into routine. We describe the architecture, the self-feeding corpus loop, the frozen-gate promotion mechanism, and an illustrative multi-domain deployment, and we discuss the economics of a classifier whose marginal labeling cost trends toward zero.
| Comments: | 9 pages, 2 figures |
| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG) |
| ACM classes: | I.2.7; I.5.4; I.2.6 |
| Cite as: | arXiv:2607.18358 [cs.CL] |
| (or arXiv:2607.18358v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.18358 arXiv-issued DOI via DataCite |
Submission history
From: Bogdan Raduta [view email]
[v1]
Mon, 20 Jul 2026 12:38:50 UTC (92 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.