“Are Multimodal Signals Synchronous?”: Temporal Relation of Declarative Gestures and Language Instructions in Human Robot Interaction
The result's identifiers
Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F68407700%3A21730%2F25%3A00388337" target="_blank" >RIV/68407700:21730/25:00388337 - isvavai.cz</a>
Result on the web
<a href="http://dx.doi.org/10.1109/ICDL63968.2025.11204382" target="_blank" >http://dx.doi.org/10.1109/ICDL63968.2025.11204382</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.1109/ICDL63968.2025.11204382" target="_blank" >10.1109/ICDL63968.2025.11204382</a>
Alternative languages
Result language
angličtina
Original language name
“Are Multimodal Signals Synchronous?”: Temporal Relation of Declarative Gestures and Language Instructions in Human Robot Interaction
Original language description
Human communication consists of multimodal signals, such as speech and gestures. These signals are not always aligned, making it difficult for artificial systems (e.g. robots) to find the proper mapping between a particular gesture and the corresponding part of a spoken instruction. The goal of our study is to identify whether and how declarative gestures during human-robot interaction are temporally synchronized with specific segments of language instructions. We conducted an experiment focused on this phenomenon, in which 26 participants taught a humanoid robot using declarative gestures and verbal instructions. The experiment was carried out in a virtual reality (VR) environment that allowed a precise capture of human movements. Gesture trajectories were annotated for the onset, peak and offset events, and statistically compared with the onset, matching language part, and offset of the instruction. The results indicate that there are significant differences between the speech and gesture onset times (W=348,495, p<0.001), with an average temporal difference of 0.56 ± 1.30 seconds, as well as between the gesture peaks and the matching language peaks (W=287,006,p<0.001), with an average difference of 0.66±1.25 seconds. Furthermore, the total duration of both signals differs significantly (W=48,672,p<0.001, with gestures lasting longer than speech. The analysis of the data distributions further revealed that even though both signals differ in absolute timing, there exists a correlation between specific key points: onset: 0.644 (p<0.001), peak: 0.646(p<0.001). These findings suggest that humans synchronize gestures with language instructions at a relational level rather than an absolute level. The findings can be applied to the design of multimodal interfaces for humanoid robots, which should help them understand better human instructions.
Czech name
—
Czech description
—
Classification
Type
D - Article in proceedings
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Result continuities
Project
<a href="/en/project/GF23-04080L" target="_blank" >GF23-04080L: Intuitive Collaboration with Household Robots in Everyday Settings</a><br>
Continuities
P - Projekt vyzkumu a vyvoje financovany z verejnych zdroju (s odkazem do CEP)
Others
Publication year
2025
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Data specific for result type
Article name in the collection
2025 IEEE International Conference on Development and Learning (ICDL)
ISBN
979-8-3315-4343-3
ISSN
—
e-ISSN
—
Number of pages
6
Pages from-to
1-6
Publisher name
IEEE Conference Publications
Place of publication
Piscataway
Event location
Praha
Event date
Sep 16, 2025
Type of event by nationality
WRD - Celosvětová akce
UT code for WoS article
—