Automatic Extraction of Clausal Embedding Based on Large-Scale English Text Data
Identifikátory výsledku
Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F26%3AWDMXHBER" target="_blank" >RIV/00216208:11320/26:WDMXHBER - isvavai.cz</a>
Výsledek na webu
<a href="https://arxiv.org/abs/2506.14064" target="_blank" >https://arxiv.org/abs/2506.14064</a>
DOI - Digital Object Identifier
<a href="http://dx.doi.org/10.48550/arXiv.2506.14064" target="_blank" >10.48550/arXiv.2506.14064</a>
Alternativní jazyky
Jazyk výsledku
angličtina
Název v původním jazyce
Automatic Extraction of Clausal Embedding Based on Large-Scale English Text Data
Popis výsledku v původním jazyce
For linguists, embedded clauses have been of special interest because of their intricate distribution of syntactic and semantic features. Yet, current research relies on schematically created language examples to investigate these constructions, missing out on statistical information and naturally-occurring examples that can be gained from large language corpora. Thus, we present a methodological approach for detecting and annotating naturally-occurring examples of English embedded clauses in large-scale text data using constituency parsing and a set of parsing heuristics. Our tool has been evaluated on our dataset Golden Embedded Clause Set (GECS), which includes hand-annotated examples of naturally-occurring English embedded clause sentences. Finally, we present a large-scale dataset of naturally-occurring English embedded clauses which we have extracted from the open-source corpus Dolma using our extraction tool.
Název v anglickém jazyce
Automatic Extraction of Clausal Embedding Based on Large-Scale English Text Data
Popis výsledku anglicky
For linguists, embedded clauses have been of special interest because of their intricate distribution of syntactic and semantic features. Yet, current research relies on schematically created language examples to investigate these constructions, missing out on statistical information and naturally-occurring examples that can be gained from large language corpora. Thus, we present a methodological approach for detecting and annotating naturally-occurring examples of English embedded clauses in large-scale text data using constituency parsing and a set of parsing heuristics. Our tool has been evaluated on our dataset Golden Embedded Clause Set (GECS), which includes hand-annotated examples of naturally-occurring English embedded clause sentences. Finally, we present a large-scale dataset of naturally-occurring English embedded clauses which we have extracted from the open-source corpus Dolma using our extraction tool.
Klasifikace
Druh
J<sub>ost</sub> - Ostatní články v recenzovaných periodicích
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)
Návaznosti výsledku
Projekt
—
Návaznosti
—
Ostatní
Rok uplatnění
2025
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů
Údaje specifické pro druh výsledku
Název periodika
arXiv preprint arXiv:2506.14064
ISSN
2331-8422
e-ISSN
—
Svazek periodika
2025
Číslo periodika v rámci svazku
2025
Stát vydavatele periodika
US - Spojené státy americké
Počet stran výsledku
11
Strana od-do
1-11
Kód UT WoS článku
—
EID výsledku v databázi Scopus
—