Charting the Growth of Social-Physical HRI (spHRI): A Systematic Review Pipeline Augmented by Small Language Models
Quick Answer
This paper shows that A systematic review of social-physical human-robot interaction (spHRI) reveals that small language models (SLMs) can significantly enhance literature screening efficiency, identifying 39 papers overlooked by human reviewers, thus supporting scalable review practices.
Key Points
- SLMs with less than 1.5B parameters screened papers significantly faster than human reviewers.
- The combined SLM ensemble identified 39 additional relevant papers, 10.29% of the dataset.
- Results indicate SLMs can augment expert reviewers, making literature reviews more sustainable.
- Fragmented terminology in spHRI complicates systematic synthesis across various fields.
- The study highlights the potential of SLMs in enhancing large-scale review practices.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Social-physical human-robot interaction (spHRI) has grown rapidly across robotics, human-computer interaction, human-robot interaction, and haptics. Yet, fragmented terminology and inconsistent methodologies make systematic synthesis difficult. To support scalable review practices, we evaluated the extent to which small language models (SLMs; < 1.5B parameters) can assist with title and abstract screening for a large spHRI systematic review. While no SLMs matched human reviewers' performance, the models operated locally and screened papers orders of magnitude faster. The combined SLM ensemble identified 39 papers reviewers missed, representing 10.29% of the final relevant dataset. These results demonstrate that SLMs can augment, rather than replace, expert reviewers and make large-scale literature reviews accessible and sustainable.
| Comments: | 5 pages, 3 figures, 2 tables, Companion Proceedings of the 21st ACM/IEEE International Conference on Human-Robot Interaction |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Digital Libraries (cs.DL); Human-Computer Interaction (cs.HC); Robotics (cs.RO) |
| Cite as: | arXiv:2606.26382 [cs.CL] |
| (or arXiv:2606.26382v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2606.26382 arXiv-issued DOI via DataCite (pending registration) |
|
| Related DOI: | https://doi.org/10.1145/3776734.3794506
DOI(s) linking to related resources |
Submission history
From: Alexis E. Block [view email]
[v1]
Wed, 24 Jun 2026 21:09:20 UTC (1,757 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.