Diachronic Parsing of Pre-Standard Irish

Identifikátory výsledku

Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F22%3ABL4D9R7X" target="_blank" >RIV/00216208:11320/22:BL4D9R7X - isvavai.cz</a>
Výsledek na webu
<a href="https://aclanthology.org/2022.cltw-1.2" target="_blank" >https://aclanthology.org/2022.cltw-1.2</a>
DOI - Digital Object Identifier
—

Alternativní jazyky

Jazyk výsledku
angličtina
Název v původním jazyce
Diachronic Parsing of Pre-Standard Irish
Popis výsledku v původním jazyce
Irish underwent a major spelling standardization in the 1940's and 1950's, and as a result it can be challenging to apply language technologies designed for the modern language to older, “pre-standard” texts. Lemmatization, tagging, and parsing of these pre-standard texts play an important role in a number of applications, including the lexicographical work on Foclóir Stairiúil na Gaeilge, a historical dictionary of Irish covering the period from 1600 to the present. We have two main goals in this paper. First, we introduce a small benchmark corpus containing just over 3800 words, annotated according to the Universal Dependencies guidelines and covering a range of dialects and time periods since 1600. Second, we establish baselines for lemmatization, tagging, and dependency parsing on this corpus by experimenting with a variety of machine learning approaches.
Název v anglickém jazyce
Diachronic Parsing of Pre-Standard Irish
Popis výsledku anglicky
Irish underwent a major spelling standardization in the 1940's and 1950's, and as a result it can be challenging to apply language technologies designed for the modern language to older, “pre-standard” texts. Lemmatization, tagging, and parsing of these pre-standard texts play an important role in a number of applications, including the lexicographical work on Foclóir Stairiúil na Gaeilge, a historical dictionary of Irish covering the period from 1600 to the present. We have two main goals in this paper. First, we introduce a small benchmark corpus containing just over 3800 words, annotated according to the Universal Dependencies guidelines and covering a range of dialects and time periods since 1600. Second, we establish baselines for lemmatization, tagging, and dependency parsing on this corpus by experimenting with a variety of machine learning approaches.

Klasifikace

Druh
D - Stať ve sborníku
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Návaznosti výsledku

Projekt
—
Návaznosti
—

Ostatní

Rok uplatnění
2022
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Údaje specifické pro druh výsledku

Název statě ve sborníku
Proceedings of the 4th Celtic Language Technology Workshop within LREC2022
ISBN
979-10-95546-73-3
ISSN
—
e-ISSN
—
Počet stran výsledku
7
Strana od-do
7-13
Název nakladatele
European Language Resources Association
Místo vydání
—
Místo konání akce
Marseille, France
Datum konání akce
1. 1. 2022
Typ akce podle státní příslušnosti
WRD - Celosvětová akce
Kód UT WoS článku
—

Podobné výsledky(10)

Czech Text Processing with Contextual Embeddings: POS Tagging, Lemmatization, Parsing and NER In-depth evaluation of Romanian natural language processing pipelines Nefnir: A high accuracy lemmatizer for Icelandic

Co hledáte?

Rychlé hledání

Chytré vyhledávání

Diachronic Parsing of Pre-Standard Irish

Identifikátory výsledku

Alternativní jazyky

Klasifikace

Návaznosti výsledku

Ostatní

Údaje specifické pro druh výsledku

Podobné výsledky(10)

Co hledáte?

Rychlé hledání

Chytré vyhledávání

Popis výsledku

Identifikátory výsledku

Identifikátory výsledku

Alternativní jazyky

Alternativní jazyky

Klasifikace

Klasifikace

Návaznosti výsledku

Návaznosti výsledku

Ostatní

Ostatní

Údaje specifické pro druh výsledku

Údaje specifické pro druh výsledku

Podobné výsledky(10)