Building a Part-of-Speech Tagged Corpus for Drenjongke (Bhutia)

Kód výsledku v IS VaVaI
<a href="https://www.isvavai.cz/riv?ss=detail&h=RIV%2F00216208%3A11320%2F20%3A10426992" target="_blank" >RIV/00216208:11320/20:10426992 - isvavai.cz</a>
Výsledek na webu
<a href="https://www.aclweb.org/anthology/2020.aacl-srw.9" target="_blank" >https://www.aclweb.org/anthology/2020.aacl-srw.9</a>
DOI - Digital Object Identifier
—

Jazyk výsledku
angličtina
Název v původním jazyce
Building a Part-of-Speech Tagged Corpus for Drenjongke (Bhutia)
Popis výsledku v původním jazyce
This research paper reports on the generation of the first Drenjongke corpus based on texts taken from a phrase book for beginners, written in the Tibetan script. A corpus of sentences was created after correcting errors in the text scanned through optical character reading (OCR). A total of 34 Part-of-Speech (PoS) tags were defined based on manual annotation performed by the three authors, one of whom is a native speaker of Drenjongke. The first corpus of the Drenjongke language comprises 275 sentences and 1379 tokens, which we plan to expand with other materials to promote further studies of this language.
Název v anglickém jazyce
Building a Part-of-Speech Tagged Corpus for Drenjongke (Bhutia)
Popis výsledku anglicky
This research paper reports on the generation of the first Drenjongke corpus based on texts taken from a phrase book for beginners, written in the Tibetan script. A corpus of sentences was created after correcting errors in the text scanned through optical character reading (OCR). A total of 34 Part-of-Speech (PoS) tags were defined based on manual annotation performed by the three authors, one of whom is a native speaker of Drenjongke. The first corpus of the Drenjongke language comprises 275 sentences and 1379 tokens, which we plan to expand with other materials to promote further studies of this language.

Druh
O - Ostatní výsledky
CEP obor
—
OECD FORD obor
10201 - Computer sciences, information science, bioinformathics (hardware development to be 2.2, social aspect to be 5.8)

Rok uplatnění
2020
Kód důvěrnosti údajů
S - Úplné a pravdivé údaje o projektu nepodléhají ochraně podle zvláštních právních předpisů

Podobné výsledky(10)