Enhancing gaze estimation using dilated CNN architectures with transfer learning

In this paper, we present a gaze estimation network with an overall architecture mainly consisting of object detection, head pose detection, and gaze detection. For object detection, we detect objects from an input image and mark the head and two eyes. For head pose detection, we use the Hopenet network to estimate the direction of the head. This network uses separate loss functions for three angles and ensures they do not conflict with each other. The VGG16 convolutional layers in this network are replaced with dilated convolutions to expand the spatial extent of the neural network. Hence, we can better capture detailed information on head posture. For gaze detection, we estimate gaze direction by detecting RGB images of the eyes. We also use dilated convolutions in VGG16 for feature extraction and incorporate the head pose angle. We finally use three fully connected layers to detect the gaze direction from the images and incorporate transfer learning to improve estimation accuracy. Experimental results have demonstrated the effectiveness of the proposed approach compared to the previous methods.

Authors

Institutions

Publication Details

Journal
Journal of the Chinese Institute of Engineers
Published
2026-09-15
DOI
https://doi.org/10.1080/02533839.2026.2720029
Primary Topic
Gaze Tracking and Assistive Technology
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Enhancing gaze estimation using dilated CNN architectures with transfer learning

Chin‐Chen Chang, Ping‐Huan Kuo, Huei‐Yung Lin, Yen-Hsun Huang et al.
Journal of the Chinese Institute of Engineers
Gaze Tracking and Assistive Technology
article

Enhancing gaze estimation using dilated CNN architectures with transfer learning

Chin‐Chen Chang, Ping‐Huan Kuo, Huei‐Yung Lin, Yen-Hsun Huang, Alan Liu
article en

Abstract

In this paper, we present a gaze estimation network with an overall architecture mainly consisting of object detection, head pose detection, and gaze detection. For object detection, we detect objects from an input image and mark the head and two eyes. For head pose detection, we use the Hopenet network to estimate the direction of the head. This network uses separate loss functions for three angles and ensures they do not conflict with each other. The VGG16 convolutional layers in this network are replaced with dilated convolutions to expand the spatial extent of the neural network. Hence, we can better capture detailed information on head posture. For gaze detection, we estimate gaze direction by detecting RGB images of the eyes. We also use dilated convolutions in VGG16 for feature extraction and incorporate the head pose angle. We finally use three fully connected layers to detect the gaze direction from the images and incorporate transfer learning to improve estimation accuracy. Experimental results have demonstrated the effectiveness of the proposed approach compared to the previous methods.

Journal of the Chinese Institute of Engineers
National Taipei University of Technology (TW), National United University (TW), National Chung Cheng University (TW)
Openalex Percentile: Top 8%
Gaze Tracking and Assistive Technology
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.