Spoken Language Treebanks in Universal Dependencies: an Overview

The result's identifiers

Result code in IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F22%3AWDSDMB3G" target="_blank" >RIV/00216208:11320/22:WDSDMB3G - isvavai.cz</a>
Result on the web
<a href="https://aclanthology.org/2022.lrec-1.191.pdf" target="_blank" >https://aclanthology.org/2022.lrec-1.191.pdf</a>
DOI - Digital Object Identifier
—

Alternative languages

Result language
angličtina
Original language name
Spoken Language Treebanks in Universal Dependencies: an Overview
Original language description
Given the benefits of syntactically annotated collections of transcribed speech in spoken language research and applications, many spoken language treebanks have been developed in the last decades, with divergent annotation schemes posing important limitations to cross-resource explorations, such as comparing data across languages, grammatical frameworks, and language domains. As a consequence, there has been a growing number of spoken language treebanks adopting the Universal Dependencies (UD) annotation scheme, aimed at cross-linguistically consistent morphosyntactic annotation. In view of the non-central role of spoken language data within the scheme and with little in-domain consolidation to date, this paper presents a comparative overview of spoken language treebanks in UD to support cross-treebank data explorations on the one hand, and encourage further treebank harmonization on the other. Our results show that the spoken language treebanks differ considerably with respect to the inventory and the format of transcribed phenomena, as well as the principles adopted in their morphosyntactic annotation. This is particularly true for the dependency annotation of speech disfluencies, where conflicting data annotations suggest an underspecification of the guidelines pertaining to speech repairs in general and the reparandum dependency relation in particular.
Czech name
—
Czech description
—

Classification

Type
D - Article in proceedings
CEP classification
—
OECD FORD branch
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Result continuities

Project
—
Continuities
—

Others

Publication year
2022
Confidentiality
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Data specific for result type

Article name in the collection
Proceedings of the 13th Conference on Language Resources and Evaluation (LREC 2022)
ISBN
979-10-95546-72-6
ISSN
—
e-ISSN
—
Number of pages
9
Pages from-to
1798-1806
Publisher name
European Language Resources Association (ELRA)
Place of publication
—
Event location
Marseille, France
Event date
Jan 1, 2022
Type of event by nationality
WRD - Celosvětová akce
UT code for WoS article
—

Similar results(10)

Universal Dependencies 2.0 - CoNLL 2017 Shared Task Development and Test Data Universal Dependencies 2.7 Universal Dependencies 2.6

What are you looking for?

Quick search

Smart search

Spoken Language Treebanks in Universal Dependencies: an Overview

The result's identifiers

Alternative languages

Classification

Result continuities

Others

Data specific for result type

Similar results(10)

What are you looking for?

Quick search

Smart search

Result description

The result's identifiers

The result's identifiers

Alternative languages

Alternative languages

Classification

Classification

Result continuities

Result continuities

Others

Others

Data specific for result type

Data specific for result type

Similar results(10)