Binocular Gaze Estimation with Single Camera and Single Light Source
Quick Answer
This study introduces a gaze estimation method using a single camera and one light source, leveraging a virtual light source to estimate gaze with polynomial regression.
Quick Take
While performance is acceptable, it shows degradation compared to systems with two actual light sources, making it suitable for mobile eye-tracking applications.
Key Points
- Proposes a gaze estimation method with one camera and one light source.
- Introduces a 'virtual light source' for improved gaze tracking.
- Estimates gaze using polynomial regression based on pupil and glint distances.
- Performance is acceptable but inferior to systems with two light sources.
- Applicable for scenarios with limited hardware, like mobile devices.
Paper Resources
📖 Reader Mode
~2 min readAbstract:According to commonly consented theories, the minimum hardware requirement for gaze tracker is one camera and two light sources to realize gaze estimation with free head movements. However, in some scenarios such as eye tracking on mobile devices, it is preferable to use less components, especially light sources. We propose a gaze estimation method with one camera and one light source. A "virtual light source" is introduced, which is geometrically placed symmetrically to the real light source with respect to the camera, and generates a "virtual glint" in the acquired image. We estimate the "virtual glint" by exploiting the relationship between the distance between two pupils and two glints in the captured image, and estimate the gaze with polynomial regression assuming two light sources are available. A new normalization factor for regression method is verified, which turns out to be practical for one-glint system. The performance is proved to be acceptable, while degradation is noticed compared to system with two actual light sources.
| Comments: | Accepted for presentation at the 2019 International Conference on Video, Signal and Image Processing (VSIP 2019), Wuhan, China, October 29-31, 2019; published in VSIP '19: Proceedings of the 2019 International Conference on Video, Signal and Image Processing, pp. 10-14, ACM, 2020; 4 figures, 1 table; ACM Proceedings ISBN: 978-1-4503-7148-3 |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.05473 [cs.CV] |
| (or arXiv:2607.05473v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.05473 arXiv-issued DOI via DataCite |
|
| Journal reference: | VSIP '19: Proceedings of the 2019 International Conference on Video, Signal and Image Processing, pp. 10-14, ACM, 2020 |
| Related DOI: | https://doi.org/10.1145/3369318.3369326
DOI(s) linking to related resources |
Submission history
From: Tongbing Huang [view email]
[v1]
Mon, 6 Jul 2026 09:13:21 UTC (350 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.