AccessAI Nexus: A Multimodal Hands-Free Human–Computer Interaction System Using Facial Gestures and Voice Commands

This paper presents AccessAI Nexus, a multimodal human–computer interaction prototype designed to enable hands-free computer interaction through facial gestures and voice commands. The system combines webcam-based facial landmark processing, speech recognition, intent processing, and computer-control automation to support cursor movement, clicking, scrolling, application control, browser navigation, keyboard actions, and voice typing. The prototype was evaluated in a preliminary single-user study consisting of 65 task trials: 30 facial-interaction trials, 30 voice-command trials, and 5 voice-typing trials. Facial tasks achieved a 100% task success rate across the recorded trials, while voice-command recognition and execution achieved 93.33% and 80.00%, respectively. Voice typing was correct in 60% of the five trials. These results characterize the behavior of the tested prototype and identify execution reliability and speech-related errors as areas for improvement. The study is preliminary and does not establish clinical effectiveness or generalizability to the broader population.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-21
DOI
https://doi.org/10.5281/zenodo.22870156
Primary Topic
Gaze Tracking and Assistive Technology
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

AccessAI Nexus: A Multimodal Hands-Free Human–Computer Interaction System Using Facial Gestures and Voice Commands

Suprabha Sinha
Zenodo (CERN European Organization for Nuclear Research)
Gaze Tracking and Assistive Technology
preprint

AccessAI Nexus: A Multimodal Hands-Free Human–Computer Interaction System Using Facial Gestures and Voice Commands

Suprabha Sinha
preprint en

Abstract

This paper presents AccessAI Nexus, a multimodal human–computer interaction prototype designed to enable hands-free computer interaction through facial gestures and voice commands. The system combines webcam-based facial landmark processing, speech recognition, intent processing, and computer-control automation to support cursor movement, clicking, scrolling, application control, browser navigation, keyboard actions, and voice typing. The prototype was evaluated in a preliminary single-user study consisting of 65 task trials: 30 facial-interaction trials, 30 voice-command trials, and 5 voice-typing trials. Facial tasks achieved a 100% task success rate across the recorded trials, while voice-command recognition and execution achieved 93.33% and 80.00%, respectively. Voice typing was correct in 60% of the five trials. These results characterize the behavior of the tested prototype and identify execution reliability and speech-related errors as areas for improvement. The study is preliminary and does not establish clinical effectiveness or generalizability to the broader population.

Zenodo (CERN European Organization for Nuclear Research)
Gaze Tracking and Assistive Technology
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

AccessAI Nexus: A Multimodal Hands-Free Human–Computer Interaction System Using Facial Gestures and Voice Commands — Suprabha Sinha · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS