Classroom Behavior Monitoring with YOLO An Empirical Study in Higher Education Settings
Quick Answer
This study presents the BAV-Classroom dataset for monitoring classroom behavior using YOLOv11, which outperformed other models.
Quick Take
It reveals significant drops in student engagement towards the end of lectures, emphasizing the need for improved teaching strategies.
Key Points
- Introduced the BAV-Classroom dataset with nine behavioral categories.
- YOLOv11 achieved the highest performance among evaluated models.
- Student concentration notably decreases during the final lecture segments.
- Demonstrates the feasibility of automated classroom monitoring.
- Insights can aid in enhancing academic quality management.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Classroom behavior monitoring plays a vital role in evaluating student engagement and improving teaching effectiveness. Traditional observation methods remain subjective and lack scalability. This study introduces a real-world dataset of classroom videos collected at the Banking Academy of Vietnam (BAV-Classroom dataset), annotated with nine distinctive behavioral categories. State-of-the-art Computer Vision models were evaluated and compared, with YOLOv11 achieving the best performance. Experimental results indicate that students' concentration often decreases notably during the final part of lectures, highlighting challenges in sustaining engagement. Our findings demonstrate the feasibility of applying computer vision for automated classroom monitoring, providing valuable insights for academic quality management.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Computers and Society (cs.CY) |
| Cite as: | arXiv:2607.02580 [cs.CV] |
| (or arXiv:2607.02580v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.02580 arXiv-issued DOI via DataCite |
Submission history
From: Trong Sinh Vu [view email]
[v1]
Tue, 30 Jun 2026 23:53:55 UTC (449 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.