“Are Multimodal Signals Synchronous?”: Temporal Relation of Declarative Gestures and Language Instructions in Human Robot Interaction
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F68407700%3A21730%2F25%3A00388337" target="_blank" >RIV/68407700:21730/25:00388337 - isvavai.cz</a>
Výsledek na webu
<a href="http://dx.doi.org/10.1109/ICDL63968.2025.11204382" target="_blank" >http://dx.doi.org/10.1109/ICDL63968.2025.11204382</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1109/ICDL63968.2025.11204382" target="_blank" >10.1109/ICDL63968.2025.11204382</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
“Are Multimodal Signals Synchronous?”: Temporal Relation of Declarative Gestures and Language Instructions in Human Robot Interaction
Popis výsledku v původním jazyce
Human communication consists of multimodal signals, such as speech and gestures. These signals are not always aligned, making it difficult for artificial systems (e.g. robots) to find the proper mapping between a particular gesture and the corresponding part of a spoken instruction. The goal of our study is to identify whether and how declarative gestures during human-robot interaction are temporally synchronized with specific segments of language instructions. We conducted an experiment focused on this phenomenon, in which 26 participants taught a humanoid robot using declarative gestures and verbal instructions. The experiment was carried out in a virtual reality (VR) environment that allowed a precise capture of human movements. Gesture trajectories were annotated for the onset, peak and offset events, and statistically compared with the onset, matching language part, and offset of the instruction. The results indicate that there are significant differences between the speech and gesture onset times (W=348,495, p<0.001), with an average temporal difference of 0.56 ± 1.30 seconds, as well as between the gesture peaks and the matching language peaks (W=287,006,p<0.001), with an average difference of 0.66±1.25 seconds. Furthermore, the total duration of both signals differs significantly (W=48,672,p<0.001, with gestures lasting longer than speech. The analysis of the data distributions further revealed that even though both signals differ in absolute timing, there exists a correlation between specific key points: onset: 0.644 (p<0.001), peak: 0.646(p<0.001). These findings suggest that humans synchronize gestures with language instructions at a relational level rather than an absolute level. The findings can be applied to the design of multimodal interfaces for humanoid robots, which should help them understand better human instructions.
Název v anglickém jazyce
“Are Multimodal Signals Synchronous?”: Temporal Relation of Declarative Gestures and Language Instructions in Human Robot Interaction
Popis výsledku anglicky
Human communication consists of multimodal signals, such as speech and gestures. These signals are not always aligned, making it difficult for artificial systems (e.g. robots) to find the proper mapping between a particular gesture and the corresponding part of a spoken instruction. The goal of our study is to identify whether and how declarative gestures during human-robot interaction are temporally synchronized with specific segments of language instructions. We conducted an experiment focused on this phenomenon, in which 26 participants taught a humanoid robot using declarative gestures and verbal instructions. The experiment was carried out in a virtual reality (VR) environment that allowed a precise capture of human movements. Gesture trajectories were annotated for the onset, peak and offset events, and statistically compared with the onset, matching language part, and offset of the instruction. The results indicate that there are significant differences between the speech and gesture onset times (W=348,495, p<0.001), with an average temporal difference of 0.56 ± 1.30 seconds, as well as between the gesture peaks and the matching language peaks (W=287,006,p<0.001), with an average difference of 0.66±1.25 seconds. Furthermore, the total duration of both signals differs significantly (W=48,672,p<0.001, with gestures lasting longer than speech. The analysis of the data distributions further revealed that even though both signals differ in absolute timing, there exists a correlation between specific key points: onset: 0.644 (p<0.001), peak: 0.646(p<0.001). These findings suggest that humans synchronize gestures with language instructions at a relational level rather than an absolute level. The findings can be applied to the design of multimodal interfaces for humanoid robots, which should help them understand better human instructions.
Klasifikace
Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
<a href="/cs/project/GF23-04080L" target="_blank" >GF23-04080L: Intuitivní spolupráce s domácím robotem během každodenních úloh</a><br>
Návaznosti
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název statě ve sborníku
2025 IEEE International Conference on Development and Learning (ICDL)
ISBN
979-8-3315-4343-3
ISSN
—
e-ISSN
—
Počet stran výsledku
6
Strana od-do
1-6
Název nakladatele
IEEE Conference Publications
Místo vydání
Piscataway
Místo konání akce
Praha
Datum konání akce
16. 9. 2025
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
—