Grammar Concept Annotation at Scale: Deployed Fine-Tuned Small Language Models Outperform Prompted Frontier Models
Quick Answer
This paper shows that Fine-tuned Qwen3.5 small language models (SLMs) outperform prompted GPT-5.4 and GPT-5.6 in grammar mastery tracking, achieving 16x cost reduction and improving learner engagement by 15.8%.
Quick Take
The 0.8B model excels in precision and recall across strict benchmarks, benefiting English learners on the platform.
Key Points
- 0.8B Qwen3.5 model reduces serving costs by approximately 16x.
- Deployed model outperforms GPT-5.4 and GPT-5.6 in precision and recall.
- Learner engagement increased by 15.8% with the new model.
- Scheduled hours and GMV from new lessons rose by 2.1% and 13.2%, respectively.
- Internalized annotation contract allows efficient model deployment.
DeepSignal Analysis
What happened
Fine-tuned Qwen3.5 small language models (SLMs) have been deployed for grammar mastery tracking, outperforming prompted GPT-5.4 and GPT-5.6 models. The 0.8B model shows significant improvements in precision and recall while reducing operational costs by 16 times. Additionally, learner engagement increased by 15.8%.
Key evidence
- The 0.8B fine-tuned Qwen3.5 model outperformed both GPT-5.4 and GPT-5.6 in precision and recall on two human-curated benchmarks.
- The deployment of the 0.8B model resulted in a 16x reduction in serving costs compared to prompted frontier models.
- Learner engagement improved by 15.8%, alongside increases in scheduled hours by 2.1% and GMV from new lessons by 13.2%.
Why it matters
The deployment of the Qwen3.5 model represents a significant advancement in language learning technology, particularly in providing cost-effective and efficient grammar mastery tracking. The ability to achieve better performance at a lower cost could lead to broader adoption of AI-driven educational tools, enhancing learning outcomes for English learners. The increase in engagement and business metrics suggests that such models can have a positive impact on both educational effectiveness and commercial viability.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Corrective feedback is among the best-evidenced drivers of second-language acquisition, yet corrections delivered during lessons rarely accumulate into an actionable view of grammar mastery. Prompted frontier models can provide such a view from learner--tutor lesson transcripts, but they are costly at scale. We close this gap by fine-tuning Qwen3.5 small language models (SLMs) on filtered and rebalanced teacher-generated supervision, then deploying an efficient 0.8B model in an end-to-end grammar mastery tracker for all English learners on our platform. Internalizing the annotation contract into adapter weights enables pairing the 0.8B model with a compact matched prompt rather than verbose instructions. On two human-curated benchmarks, both the deployed 0.8B model and a 4B reference comparator outperform prompted GPT-5.4 and GPT-5.6 Sol in precision and recall under nested matching criteria of increasing strictness: concept, evidence span, and correctness. The deployed 0.8B SLM reduces serving cost by approximately 16$\times$. A feature-level online experiment shows significant gains in learner engagement ($+15.8\%$) and key business metrics, including scheduled hours ($+2.1\%$) and GMV from new lessons ($+13.2\%$).
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| ACM classes: | I.2.7; I.2.6; K.3.1 |
| Cite as: | arXiv:2610.10827 [cs.CL] |
| (or arXiv:2610.10827v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10827 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Marjan Celikik [view email]
[v1]
Wed, 7 Oct 2026 19:31:24 UTC (144 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →The "10th Juror": Open-Set Standpoint Screening for Bureaucratic Bias Detection
MARS-Gov introduces a framework for detecting bureaucratic bias in Dutch government documents, achieving a new state-of-the-art F1 score of 0.880. This model outperforms existing zero-shot detectors by 20.2 points and reduces unnecessary interventions to just 2.5%. The framework's dynamic '10th juror' adapts to emerging biases, enhancing legal language processing.