An Empirical Study of Handcrafted Feature Learning and Convolutional Neural Networks for Facial Expression Recognition
Quick Answer
This study evaluates handcrafted features like HOG and LBP against CNNs for facial expression recognition across three datasets (FER-2013, CK+, KDEF).
Quick Take
CNNs outperform traditional methods, especially on complex data, while HOG excels in controlled settings, and LBP underperforms overall, emphasizing the need for robust feature learning in real-world applications.
Key Points
- CNNs achieved the best performance overall, especially on complex datasets.
- HOG performed well in controlled environments but not in varied conditions.
- LBP consistently underperformed across all tested datasets.
- Dataset complexity significantly impacts recognition performance.
- Robust feature learning is crucial for effective real-world applications.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Facial expression recognition is an important computer vision task with applications in human--computer interaction, mental health monitoring, driver alert systems, and behavioral analysis. While convolutional neural networks (CNNs) dominate modern facial expression recognition, handcrafted feature descriptors such as Histogram of Oriented Gradients (HOG) and Local Binary Patterns (LBP) remain useful classical baselines. This study compares HOG with Support Vector Machine (SVM), LBP with Logistic Regression, and a lightweight CNN across three facial expression datasets: FER-2013, CK+, and KDEF. The results show that CNNs achieve the best overall performance, particularly on more complex data, while HOG performs strongly in controlled environments. LBP performs poorly across all datasets. The study highlights that dataset complexity significantly affects performance and that robust feature learning is essential for real-world facial expression recognition.
| Comments: | 9 pages, 14 figures, 6 tables |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG) |
| Cite as: | arXiv:2607.15288 [cs.CV] |
| (or arXiv:2607.15288v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.15288 arXiv-issued DOI via DataCite |
Submission history
From: Rallage Rangika Chethiya Bandara Galkaduwa [view email]
[v1]
Fri, 19 Jun 2026 02:01:11 UTC (4,521 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.