Can a System-One LLM Perform Knowledge Tracing When Few or No Learners Are Logged?
Quick Answer
The System-One LLM, Jev, outperforms traditional deep knowledge tracing models by achieving a mean AUC of 0.706 on seven datasets, significantly surpassing the best deep model's 0.689, and does so at a fraction of the cost.
Quick Take
With additional data from logged learners, JevKT improves to 0.722, maintaining an advantage even with as few as 16 learners logged.
Key Points
- Jev achieves a mean AUC of 0.706 without any target platform data.
- Outperforms 28 deep KT models, which had a maximum AUC of 0.689.
- JevKT improves performance to 0.722 with additional logged learner data.
- Significant advantage persists with as few as 16 learners logged.
- Deep KT models only catch up with 64 to 128 learners logged.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Knowledge tracing (KT) models need many logged learners, so a new course or platform starts without a usable model. In LLM-based KT the LLM generates the answer, which we call System-Two; it is either fine-tuned on the target data or reasons and votes over ten samples, which is slow and gives coarse probabilities. We ask whether an off-the-shelf System-One LLM, which returns a probability for a typed question directly in a single pass, can perform KT when few or no learners are logged. On seven datasets, Jev without any data from the target platform reaches a mean AUC of .706, above the best of 28 deep KT models trained on 8 learners (.689) and above System-Two Thinking-KT on all seven datasets (.650) at about 1/100 of its API cost. Adding examples and a similar-learner statistic from the logged learners (JevKT) raises this to .722; JevKT stays significantly ahead of deep KT up to 16 learners and ahead on average up to 64, and supervised KT catches up between 64 and 128 learners. Among the readers we tested, the gain is specific to Jev, since three other LLMs queried with the byte-identical typed request through the official System-One adapter fall below it on all seven datasets, and reader swaps and contamination checks find no evidence that the input format or memorised data explain the gain. For new learners the advantage holds from their first interactions, whereas on unseen items with all learners logged, deep KT remains ahead.
| Comments: | 41 pages, 7 figures. Code and result summaries: this https URL |
| Subjects: | Computation and Language (cs.CL); Computers and Society (cs.CY); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.11135 [cs.CL] |
| (or arXiv:2610.11135v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11135 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Unggi Lee [view email]
[v1]
Thu, 8 Oct 2026 02:59:20 UTC (82 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →The "10th Juror": Open-Set Standpoint Screening for Bureaucratic Bias Detection
MARS-Gov introduces a framework for detecting bureaucratic bias in Dutch government documents, achieving a new state-of-the-art F1 score of 0.880. This model outperforms existing zero-shot detectors by 20.2 points and reduces unnecessary interventions to just 2.5%. The framework's dynamic '10th juror' adapts to emerging biases, enhancing legal language processing.