Investigation of Word Recognition Methods for Japanese Fingerspelling Using Large Language Models
The development of automatic fingerspelling recognition systems is critical for facilitating seamless communication for the deaf and hard-of-hearing community. Historically, recognition technologies relied on wearable sensors, such as data gloves, inertial measurement units (IMUs), and surface electromyography (sEMG), which provided stable kinematic data but suffered from high hardware costs and user discomfort. Non-invasive vision-based sensing has emerged to eliminate wearable overhead but is subject to severe physical bottlenecks: continuous high-speed movements inevitably introduce motion blur and self-occlusion, drastically degrading tracking accuracy. To address these visual sensing constraints, we investigate a visual-language integration framework that combines dynamic visual representations derived from a single RGB video stream with the lexical knowledge of Large Language Models (LLMs), allowing the language model to compensate for uncertainty or errors in the visual representation. Our approach employs a robust visual encoder—comprising a Temporal Convolutional Network (TCN) and a Transformer—to extract and integrate dynamic transition features from 3D skeletal coordinates. Rather than propose a new LLM architecture, this study focuses on how visual information should be represented and transferred to an LLM by comparing discrete-text, probability-level, and latent-feature-level integration strategies under the same experimental conditions. Specifically, we systematically evaluated three integration strategies: Text Input (TI), Projector Probabilities Input (PPI), and Projector Transformer Features Input (PTFI). Experimental results on a 28-word closed-set evaluation show that while the word-level exact-match accuracy (WA) of the visual encoder alone was limited to 11.66% with continuous signing, our LLM-integrated architecture elevated the WA to a maximum of 87.73%. Crucially, the PTFI approach successfully suppressed character-level degradation by grounding linguistic reasoning in latent Transformer-based visual evidence. These results suggest that integrating LLM-based linguistic reasoning with dynamic visual features is a promising approach for closed-set continuous fingerspelling recognition under controlled experimental conditions.
Authors
- Ryota Murai (ORCID: https://orcid.org/0009-0007-0248-9267)
- Yousun Kang (ORCID: https://orcid.org/0009-0006-4391-7064)
- Duk Shin (ORCID: https://orcid.org/0000-0001-8454-0319)
- Naoto Tsuta
- Renta Kuzuyama
Institutions
- Tokyo Polytechnic University (JP)
Publication Details
- Journal
- Sensors
- Published
- 2026-09-24
- DOI
- https://doi.org/10.3390/s26196044
- Primary Topic
- Hand Gesture Recognition Systems
- Type
- article
- Field-Weighted Citation Impact
- 0.00