Investigation of Word Recognition Methods for Japanese Fingerspelling Using Large Language Models

The development of automatic fingerspelling recognition systems is critical for facilitating seamless communication for the deaf and hard-of-hearing community. Historically, recognition technologies relied on wearable sensors, such as data gloves, inertial measurement units (IMUs), and surface electromyography (sEMG), which provided stable kinematic data but suffered from high hardware costs and user discomfort. Non-invasive vision-based sensing has emerged to eliminate wearable overhead but is subject to severe physical bottlenecks: continuous high-speed movements inevitably introduce motion blur and self-occlusion, drastically degrading tracking accuracy. To address these visual sensing constraints, we investigate a visual-language integration framework that combines dynamic visual representations derived from a single RGB video stream with the lexical knowledge of Large Language Models (LLMs), allowing the language model to compensate for uncertainty or errors in the visual representation. Our approach employs a robust visual encoder—comprising a Temporal Convolutional Network (TCN) and a Transformer—to extract and integrate dynamic transition features from 3D skeletal coordinates. Rather than propose a new LLM architecture, this study focuses on how visual information should be represented and transferred to an LLM by comparing discrete-text, probability-level, and latent-feature-level integration strategies under the same experimental conditions. Specifically, we systematically evaluated three integration strategies: Text Input (TI), Projector Probabilities Input (PPI), and Projector Transformer Features Input (PTFI). Experimental results on a 28-word closed-set evaluation show that while the word-level exact-match accuracy (WA) of the visual encoder alone was limited to 11.66% with continuous signing, our LLM-integrated architecture elevated the WA to a maximum of 87.73%. Crucially, the PTFI approach successfully suppressed character-level degradation by grounding linguistic reasoning in latent Transformer-based visual evidence. These results suggest that integrating LLM-based linguistic reasoning with dynamic visual features is a promising approach for closed-set continuous fingerspelling recognition under controlled experimental conditions.

Authors

Institutions

Publication Details

Journal
Sensors
Published
2026-09-24
DOI
https://doi.org/10.3390/s26196044
Primary Topic
Hand Gesture Recognition Systems
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Investigation of Word Recognition Methods for Japanese Fingerspelling Using Large Language Models

Ryota Murai, Yousun Kang, Duk Shin, Naoto Tsuta et al.
Sensors
Hand Gesture Recognition Systems
article

Investigation of Word Recognition Methods for Japanese Fingerspelling Using Large Language Models

Ryota Murai, Yousun Kang, Duk Shin, Naoto Tsuta, Renta Kuzuyama
article en

Abstract

The development of automatic fingerspelling recognition systems is critical for facilitating seamless communication for the deaf and hard-of-hearing community. Historically, recognition technologies relied on wearable sensors, such as data gloves, inertial measurement units (IMUs), and surface electromyography (sEMG), which provided stable kinematic data but suffered from high hardware costs and user discomfort. Non-invasive vision-based sensing has emerged to eliminate wearable overhead but is subject to severe physical bottlenecks: continuous high-speed movements inevitably introduce motion blur and self-occlusion, drastically degrading tracking accuracy. To address these visual sensing constraints, we investigate a visual-language integration framework that combines dynamic visual representations derived from a single RGB video stream with the lexical knowledge of Large Language Models (LLMs), allowing the language model to compensate for uncertainty or errors in the visual representation. Our approach employs a robust visual encoder—comprising a Temporal Convolutional Network (TCN) and a Transformer—to extract and integrate dynamic transition features from 3D skeletal coordinates. Rather than propose a new LLM architecture, this study focuses on how visual information should be represented and transferred to an LLM by comparing discrete-text, probability-level, and latent-feature-level integration strategies under the same experimental conditions. Specifically, we systematically evaluated three integration strategies: Text Input (TI), Projector Probabilities Input (PPI), and Projector Transformer Features Input (PTFI). Experimental results on a 28-word closed-set evaluation show that while the word-level exact-match accuracy (WA) of the visual encoder alone was limited to 11.66% with continuous signing, our LLM-integrated architecture elevated the WA to a maximum of 87.73%. Crucially, the PTFI approach successfully suppressed character-level degradation by grounding linguistic reasoning in latent Transformer-based visual evidence. These results suggest that integrating LLM-based linguistic reasoning with dynamic visual features is a promising approach for closed-set continuous fingerspelling recognition under controlled experimental conditions.

SensorsVol. 26(19)
Tokyo Polytechnic University (JP)
Quality Education
Openalex Percentile: Top 9%
Hand Gesture Recognition Systems
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.