LLM-Based Waypoint Authoring via Augmented Reality for Human–Robot Interaction

Augmented reality (AR) can make robot programming more intuitive by letting users set motion parameters directly in the workspace. This paper extends AURaPath, an AR system for interactive 6-DoF waypoint determination on a UR10 collaborative robot using HoloLens2, with a voice modality for hands-free authoring. The voice layer follows a two-tier design: an on-device keyword recognizer handles the safety-critical verbs behind a two-step confirmation gate, and a large language model (LLM) maps the resulting transcript to the system’s structured command set so that operators need not memorize fixed phrases; the deployed pipeline takes spoken microphone input, while the benchmark below evaluates the transcript-to-command mapping stage in isolation. We evaluate the mapping layer on an objective benchmark of 169 labeled utterances that isolates language-to-command mapping from speech recognition and robot execution, and we compare a deterministic keyword baseline, four local models on a single consumer GPU, and free-tier cloud backends. The keyword baseline cannot express parameterized authoring commands and scores zero on every authoring field, while a free-tier cloud model reaches 95.3% exact-match results, averaged over 10 runs. The reject analysis shows that the ability to decline out-of-scope input, rather than raw accuracy, determines whether a model is suitable for the language-mapping tier of our safety-gated architecture, in which a deterministic layer and a preview-before-execute gate remain responsible for execution, which supports keeping the safety-critical verbs on the deterministic layer. The benchmark selects the backend that the deployed system uses, on the same physical UR10 demonstrated in the base system. The project code, the evaluation harness and dataset, and the base-system demonstration are available on GitHub (Please see Data Availability Statement section).

Authors

Institutions

Publication Details

Journal
Applied Sciences
Published
2026-10-09
DOI
https://doi.org/10.3390/app16209974
Primary Topic
Augmented Reality Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

LLM-Based Waypoint Authoring via Augmented Reality for Human–Robot Interaction

Saeed Mozaffari, Shahpour Alirezaee, Kasra Mojallal, Mohammad Hashemi
Applied Sciences
Augmented Reality Applications
article

LLM-Based Waypoint Authoring via Augmented Reality for Human–Robot Interaction

Saeed Mozaffari, Shahpour Alirezaee, Kasra Mojallal, Mohammad Hashemi
article en

Abstract

Augmented reality (AR) can make robot programming more intuitive by letting users set motion parameters directly in the workspace. This paper extends AURaPath, an AR system for interactive 6-DoF waypoint determination on a UR10 collaborative robot using HoloLens2, with a voice modality for hands-free authoring. The voice layer follows a two-tier design: an on-device keyword recognizer handles the safety-critical verbs behind a two-step confirmation gate, and a large language model (LLM) maps the resulting transcript to the system’s structured command set so that operators need not memorize fixed phrases; the deployed pipeline takes spoken microphone input, while the benchmark below evaluates the transcript-to-command mapping stage in isolation. We evaluate the mapping layer on an objective benchmark of 169 labeled utterances that isolates language-to-command mapping from speech recognition and robot execution, and we compare a deterministic keyword baseline, four local models on a single consumer GPU, and free-tier cloud backends. The keyword baseline cannot express parameterized authoring commands and scores zero on every authoring field, while a free-tier cloud model reaches 95.3% exact-match results, averaged over 10 runs. The reject analysis shows that the ability to decline out-of-scope input, rather than raw accuracy, determines whether a model is suitable for the language-mapping tier of our safety-gated architecture, in which a deterministic layer and a preview-before-execute gate remain responsible for execution, which supports keeping the safety-critical verbs on the deterministic layer. The benchmark selects the backend that the deployed system uses, on the same physical UR10 demonstrated in the base system. The project code, the evaluation harness and dataset, and the base-system demonstration are available on GitHub (Please see Data Availability Statement section).

Applied SciencesVol. 16(20)
University of Windsor (CA)
Openalex Percentile: Top 15%
Augmented Reality Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

LLM-Based Waypoint Authoring via Augmented Reality for Human–Robot Interaction — Saeed Mozaffari, Shahpour Alirezaee, et al. · Applied Sciences (2026) | TGRS Research Map | TGRS