CSTutorBench: Benchmarking Small Language Models as Tutors for Block-Based Programming
Quick Answer
CSTutorBench introduces a benchmark for evaluating small language models (SLMs) as tutors in block-based programming, revealing that while models excel in vocabulary and tone, they struggle with deeper pedagogical behaviors.
Quick Take
The study indicates that model family and instruction-tuning are better predictors of tutoring quality than parameter count alone, with targeted prompt revisions improving performance for 10 out of 11 models tested.
Key Points
- CSTutorBench evaluates 11 models ranging from 4B to 120B parameters.
- Models perform well on vocabulary but struggle with pedagogical engagement.
- Model family and instruction-tuning are better quality predictors than size.
- Targeted prompt revisions improved scores for 10 out of 11 models.
- Benchmark focuses on block-based programming in VEX VR environment.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Large language models are increasingly explored as AI tutors, yet deploying them in K-12 settings raises concerns around privacy, cost, and reliance on proprietary models. Small language models (SLMs) offer a promising alternative, but selecting the right model for a specific educational context remains difficult, particularly when the target domain, such as block-based programming, is largely absent from model training data. We introduce CSTutorBench, a benchmark for evaluating language models as CS tutors in VEX VR, a block-based robotics environment. The benchmark comprises 17 scenario-based questions scored against a pedagogical rubric grounded in established tutoring and feedback research, with a human-in-the-loop LLM-as-judge pipeline for evaluation. Preliminary findings across 11 models (4B-120B parameters) reveal that models perform well on surface-level criteria such as vocabulary and tone but struggle with deeper pedagogical behaviors, particularly avoiding answer leakage and engaging with student debugging histories. In our sample, model family and instruction-tuning approach appear to be better predictors of tutoring quality than parameter count alone, though the small number of models limits the strength of this conclusion. A targeted prompt revision grounded in recent educational prompt engineering research improved scores for 10 of 11 models. These results underscore the value of context-specific, pedagogically grounded benchmarks for SLM selection in educational deployment.
| Subjects: | Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC) |
| Cite as: | arXiv:2607.05571 [cs.AI] |
| (or arXiv:2607.05571v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.05571 arXiv-issued DOI via DataCite (pending registration) |
|
| Journal reference: | SLM4ED'26: The 1st Workshop of Small Language Models for Education (SLM4ED). AIED 2026. Seoul, Republic of Korea |
Submission history
From: H Chad Lane [view email]
[v1]
Mon, 6 Jul 2026 19:15:07 UTC (1,440 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.