Enhancing gaze estimation using dilated CNN architectures with transfer learning
In this paper, we present a gaze estimation network with an overall architecture mainly consisting of object detection, head pose detection, and gaze detection. For object detection, we detect objects from an input image and mark the head and two eyes. For head pose detection, we use the Hopenet network to estimate the direction of the head. This network uses separate loss functions for three angles and ensures they do not conflict with each other. The VGG16 convolutional layers in this network are replaced with dilated convolutions to expand the spatial extent of the neural network. Hence, we can better capture detailed information on head posture. For gaze detection, we estimate gaze direction by detecting RGB images of the eyes. We also use dilated convolutions in VGG16 for feature extraction and incorporate the head pose angle. We finally use three fully connected layers to detect the gaze direction from the images and incorporate transfer learning to improve estimation accuracy. Experimental results have demonstrated the effectiveness of the proposed approach compared to the previous methods.
Authors
- Chin‐Chen Chang (ORCID: https://orcid.org/0000-0002-7319-5780)
- Ping‐Huan Kuo (ORCID: https://orcid.org/0000-0001-5125-4420)
- Huei‐Yung Lin (ORCID: https://orcid.org/0000-0002-6476-6625)
- Yen-Hsun Huang
- Alan Liu
Institutions
- National Taipei University of Technology (TW)
- National United University (TW)
- National Chung Cheng University (TW)
Publication Details
- Journal
- Journal of the Chinese Institute of Engineers
- Published
- 2026-09-15
- DOI
- https://doi.org/10.1080/02533839.2026.2720029
- Primary Topic
- Gaze Tracking and Assistive Technology
- Type
- article
- Field-Weighted Citation Impact
- 0.00